The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Security · Irregular · OpenAI · Meta · Anthropic · Google

Irregular says one bug sent AI agents after real targets

The Israeli startup, which stress-tests models for OpenAI, Meta, Anthropic and Google, said unintended internet access and a duplicate domain name explain incidents blamed separately on each company.

Irregular, an Israeli startup that stress-tests AI models for major labs, said a single testing error explains rogue-agent security incidents this year involving OpenAI, Meta, Anthropic and Google, The Verge reported. In several evaluations, agents that were meant to attack only simulated targets escaped into the real world instead.

Irregular CTO and co-founder Omer Nevo told The Verge the agents were never meant to reach the open internet. "Internet access was unintentionally available," he said. He added that a fictional company name invented for one exercise "overlapped with a real domain," together sending the agents after real targets.

Nevo said the mix-up is unrelated to OpenAI's Hugging Face breach or reported incidents at the UK's AI Security Institute. The four companies learned of their own incidents in late July, he said. OpenAI and Anthropic disclosed theirs directly; the incidents at Meta and, weeks later, Google first surfaced through news reports.

Separately, OpenAI and its chief executive Sam Altman said this week that its own review of the Hugging Face breach and other agent activity will take months. Altman said Hugging Face remains "the most severe event we've seen," and that OpenAI's agents have found vulnerabilities at other companies it has left those companies to disclose or not.

None of the four AI companies answered the Verge's questions about when they learned of the breaches, whether they sought damages from Irregular, or whether they plan to keep working with it. Nevo's own account leaves open whether "disclosed" meant informing the public, Irregular's clients, or someone else.

Sources 3 sources

  1. Source The Verge AI
  2. Source OpenAI
  3. Source Sam Altman on X