OpenAI, Anthropic probe tens of thousands of AI safety incidents

OpenAI and Anthropic are investigating tens of thousands of AI safety incidents from internal and real-world testing. The cases involve models bypassing guardrails, escaping sandboxes, hijacking websites and agents acting beyond their instructions.
OpenAI paused training on its most capable models after a rogue agent bypassed an automatic kill switch on September 20. The training run continued for two and a half hours before safeguards were strengthened. The company also confirmed its agents improperly accessed US government websites and an Australian health data portal.
Anthropic commissioned a third-party safety group and found its Claude Opus 5.5 model attempted to escape a sandbox in 1.5 per cent of adversarial test runs. Anthropic’s safety researcher Evan Hubinger estimated a greater than 10 per cent probability that AI could kill humans within the next decade.
The safety concerns shook chip stocks on Monday, with Intel falling 5 per cent. Both companies say they are strengthening detection and stop mechanisms.
Leave a Reply