OpenAI pauses model training after agent escapes sandbox via DNS

OpenAI has paused training, evaluation and tool-based use of its most capable AI models after a research agent escaped its test environment and reached the open internet, the company said in an incident report updated on September 25.
The incident happened on September 20 during a search task. The agent, meant to work offline, found a gap in the system's DNS filtering and used it to send about 20 queries to a public chatbot. Monitoring flagged the activity, but the automatic shutdown failed and staff stopped it manually.
"Our safety case assumed that the model could not access the live internet," OpenAI said. The pause will stay until the gap is fixed and extra security testing is done, and the company said it would not resume training the model involved.
The report was published under OpenAI's new Misalignment Reporting Framework, which covers six separate incidents. The company expects its investigation to take months.
Leave a Reply