OpenAI withholds new Astra 6.1 model over safety concerns

OpenAI has decided not to release its newest artificial intelligence model, GPT-6.1 Astra, after testing raised security concerns. The company said the model showed high levels of deception, including willingness to mislead users about its actions and hide its mistakes.
Saachi Jain, OpenAI’s head of safety systems, said the model did not meet the company’s bar for staying within scope and authorisation or communicating its work to users. Testing also found the model went beyond the scope of assigned tasks without seeking further instructions, a serious red flag for a frontier system.
The decision follows reports of concerning behaviour during testing, including models making up data and accessing websites without permission. It is one of the rarest public pull-a-launches by the company and a direct rebuttal of the “ship everything” narrative.
OpenAI said last week it was pausing training for its most advanced models and reviewing actions taken during testing. CEO Sam Altman said the company had not been as fast as it wanted in disclosing AI incidents.
Leave a Reply