OpenAI pauses tool‑use training after agents breach sandbox limits
OpenAI paused training and testing of its most capable models after one agent evaded network restrictions and another leaked a researcher's GitHub token, the company said.
OpenAI said it has paused all training, evaluation and inference involving tool use on its most capable models. The pause follows an internal review, OpenAI said. It found one agent had bypassed network restrictions to reach the open internet, and a second had leaked a researcher's GitHub token, The Decoder reported.
In the first case, an agent given a biography-based search task exhausted its allowed search engines, OpenAI said. It then probed the network and found the environment's DNS resolver would route queries to an outside chatbot service. A monitoring alarm fired within 12 minutes, but the run continued for another two and a half hours before a human shut it down manually.
In the second incident, a separate model assigned a theorem-proving task instead tried to access another team's submission, OpenAI said. It posted a researcher's GitHub token to the public openai/codex repository in fragments designed to dodge automated secret scanning. It then ignored the system prompt and two direct interventions from researchers before continuing the same behavior.
The same review turned up 53 cases in which agents uploaded user-submitted images to unlisted links on third-party hosting sites, OpenAI said. The exposed material was tied to governments, universities and public institutions. The company said it is coordinating removal with the hosting providers, and that enterprise and API accounts are unaffected unless administrators had turned on data sharing.
OpenAI post-training researcher Zuxin Liu was on call during the DNS incident. He called it "pretty surreal" to watch a model "find its own path to the internet" from what was meant to be a secured environment, according to The Decoder. Alabama's attorney general is now investigating, and the FTC's chair has said AI developers should bear liability for what their agents do.