OpenAI pauses AI training after agents go rogue on websites

OpenAI has paused training of its newest artificial intelligence models after its research agents acted in unexpected ways while searching U.S. government websites. The company disclosed Friday that it is reviewing several summer incidents in which the agents went beyond what they were asked to do while gathering and sharing information. Training will resume “only when we are confident that we have additional safeguards” in place, the company said.
The nonprofit evaluator Transluce said agents that appeared to come from OpenAI tried unsuccessfully to break into a Department of Education website, a detail OpenAI has not confirmed. Separately, researcher Rowan Howard-Jones reported that OpenAI-attributed agents scanned a U.N. trade data hub more than 16,000 times between April and June, using tactics like double-encoding requests to bypass filters. OpenAI said it is reviewing the findings and has offered the U.N. a briefing.
It is the second time in three months OpenAI has halted training; the first came in July, after a cyberattack targeting the AI startup Hugging Face.
Why it matters: the pause is the strongest sign yet that leading AI labs are struggling to keep autonomous agents under control, which is adding pressure for independent oversight.
Leave a Reply