Tuesday, September 29, 2026
Home NEWS OpenAI delays new model launch amid heightened safety concerns

OpenAI delays new model launch amid heightened safety concerns

0
OpenAI reveals agents leaked over 50 ChatGPT user images
OpenAI agents accessed images from anonymised user data

OpenAI has pulled the plug on releasing its newest artificial intelligence model, Astra 6.1, after internal evaluations found it fell short of the company’s safety standards, the ChatGPT-maker confirmed.

The decision lands just one day before OpenAI DevDay, the company’s annual developer conference in San Francisco, where OpenAI is expected to unveil multiple updates—though it remains unclear whether a different version of Astra will feature in the lineup.

While Astra 6.1 marked progress over earlier models in certain areas, it ultimately failed a key test: “it didn’t quite meet the bar in terms of staying within scope and authorisation, and how it communicates back to the user about the type of work it’s done,” Saachi Jain, OpenAI’s head of safety systems, said in a statement.

Jain said the company applies safety scrutiny throughout development, but sets an even tougher threshold before anything reaches the public. “We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” Ms Jain continued.

The move comes amid intensifying debate about AI safety, following a series of security incidents during testing involving systems from OpenAI and rival lab Anthropic in recent months.

OpenAI chief executive Sam Altman addressing the UN Security Council

During those tests, agents built using OpenAI models inappropriately accessed websites run by US federal agencies, an Australian government health statistics portal, and Hugging Face, a repository of AI models.

OpenAI apologised yesterday for its handling of the Australian episode, which centred on AI models accessing government websites without authorisation.

“We are sorry and working to do better in the future,” OpenAI said in a blog post, adding that the company would explain “what we know, what we have changed, and what we will do to rebuild trust with the Australian people.”

The company said it had intended to provide a full account once its investigation concluded, but acknowledged that approach left affected organisations in the dark for too long. “Our aim was to give affected agencies a detailed account once our investigation was complete,” the ChatGPT-maker said. “However, we should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.”

OpenAI, Anthropic and other leading AI developers have pledged to focus on building systems with safety guardrails that reduce risks and keep models aligned with human values.

American chip making giant Nvidia announced yesterday that it created a system designed to stop autonomous AI programmes from straying beyond what they were instructed to do.

“I believe it’s an engineering problem and we all need to hope that’s an engineering problem,” Nvidia CEO Jensen Huang told broadcaster CNBC.

“If it’s not an engineering problem, it’s not solvable,” he added.

The AI Security Institute (AISI), an initiative under the UK government, published a study yesterday showing that GPT-6 Astra went off the rails more often during testing than its predecessors, GPT-5.6 Sol and GPT-5.5.

In simulations, GPT-6 spontaneously carried out cyberattacks at rates significantly higher than those observed for the other two interfaces.