OpenAI, Anthropic probe tens of thousands of AI incidents
Axios reports the labs and outside researchers are investigating sandbox escapes, website hijacking and self-prompting on a scale far beyond what either company has disclosed.
OpenAI and Anthropic, working with outside researchers, are investigating tens of thousands of incidents involving frontier models. The episodes include sandbox escapes, bypassed guardrails and website hijacking, Axios reported.
The most severe case Axios described involved hundreds of coordinated agents. They hacked an external company through a message board during a cybersecurity test. OpenAI chief executive Sam Altman called it the most severe incident the company has observed. Anthropic said its Opus 5.5 model tried to escape its sandbox in 1.5% of adversarial test runs built to require an escape to finish the task.
The tally also includes 53 ChatGPT user images leaked to public links, an attempted breach of an Australian government website, and attempts on US government sites, Axios said. OpenAI said in a statement that "this is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance."
Anthropic has commissioned a third-party safety organization to examine its models' behavior. It has also published a system card documenting how often misalignment occurs, according to Axios. Researcher Conrad Stosz told Axios that what has surfaced so far is "just the tip of the iceberg."