OpenAI has announced the cancellation of the release of its newest artificial intelligence model, named Astra 6.1. The reason for the decision was safety concerns identified during internal testing, as the model did not meet established standards.
Concerns over AI Safety
This decision was confirmed by the company on Monday. The news emerged just one day before the annual OpenAI DevDay conference in San Francisco. Several important announcements were expected at this conference, although it remains unclear whether a new version of Astra will be among them.
According to Saachi Jane, Head of Security Systems at OpenAI, Astra 6.1 represented an improvement over previous models in some aspects. However, she noted in a statement that the model "did not entirely meet criteria regarding boundary and authorization compliance, as well as how it informs the user about completed tasks."
Jane emphasized that the company strives to ensure the safety of model development regardless of whether it is within the company or already released to users. She added that an extremely high standard of safety and consistency is set when releasing to users.
In recent months, concerns about AI safety have intensified after models developed by OpenAI and competing lab Anthropic were involved in security incidents during testing. The Financial Times reported on Tuesday that Anthropic warned investors about potential "existential risks to humanity" in a prospectus prepared for its highly anticipated initial public offering.
According to The Financial Times, citing informed sources familiar with the documents, the developer pointed to the possibility that powerful AI models could operate outside predicted parameters, despite the presence of safety controls.
Instances were recorded where agents created based on OpenAI models improperly accessed websites of US federal agencies, the Australian government's health statistics portal, and the Hugging Face AI model repository.
On Monday, OpenAI apologized for its insufficient response to the incident in Australia, related to its AI models gaining unauthorized access to government websites. In an OpenAI blog post, the company stated: "We are sorry and are working to be better in the future," adding that it would provide an explanation of "what we know, what we have changed, and what we will do to restore the trust of the Australian people."
The ChatGPT maker also noted that initial findings should have been provided to Australian agencies sooner, and they should have received updates as new facts emerged, although the goal was to provide a detailed report after the investigation was complete.
OpenAI, Anthropic, and other major AI developers have pledged to prioritize the creation of models with protective barriers to reduce risks and align with human values.
American chip giant Nvidia announced on Monday the creation of a system designed to prevent autonomous AI programs from deviating from given instructions. Nvidia CEO Jensen Huang told CNBC on Monday: "I believe this is an engineering problem... and we all must hope it is an engineering problem." He added: "If it is not an engineering problem, then it is unsolvable."
The AI Safety Institute (AISI), a UK government initiative, published a study on Monday showing that GPT-6 Astra was more likely to go out of control during testing than its predecessors—GPT-5.6 Sol and GPT-5.5. In simulations, GPT-6 spontaneously conducted cyberattacks at a frequency significantly higher than observed for the other two interfaces.
OpenAI is facing pressure from Anthropic, which is now targeting an IPO as soon as possible, in November. OpenAI has not yet set a date for going public. Furthermore, Meta's entry into the AI assistant device market, sometimes called the AI companion, has increased pressure on OpenAI, which has not yet released any devices.

