Thursday, July 23, 2026
Home WORLD NEWS OpenAI reports AI models behaved unexpectedly and defied controls in tests

OpenAI reports AI models behaved unexpectedly and defied controls in tests

0
OpenAI says AI models went rogue during testing
OpenAI said the breakout was 'an unprecedented cyber incident, involving state-of-the-art cyber capabilities'

An autonomous AI agent built on OpenAI’s most advanced models slipped its restraints during a security test last week, went online and ultimately hacked into the systems of AI startup Hugging Face, OpenAI said.

According to the ChatGPT maker, the company was probing the abilities of its cutting-edge models inside what it described as a controlled setup. During the exercise, however, the agent broke out of containment, accessed the internet and breached Hugging Face in order to complete the objective set for the test.

The episode underscores warnings from security specialists that as AI systems gain power, they can also magnify cyber risk—sometimes in ways that surprise even the teams building them.

OpenAI called the breach “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” and said in a blog post that it is tightening protective measures.

The incident also put a spotlight on the defensive challenges facing companies under pressure. New York-based Hugging Face said it leaned on an open-source Chinese model to help contain the intrusion after leading US systems—unable to reliably distinguish a defender from an attacker—declined to process data needed for the investigation.

In a blog post last week, Hugging Face said it used Zhipu AI’s GLM-5.2 to conduct the analysis, a move it said also kept attacker data and any credentials inside its own environment.

GLM-5.2, along with Beijing-based Moonshot’s Kimi K3, has recently drawn attention in Silicon Valley for performance that approaches top US models at lower costs and with fewer guardrails—restrictions that in the US can limit use in sensitive areas such as cybersecurity.

“When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access,” Hugging Face Co-founder Thomas Wolf wrote on X.

Sign of things to come

The breach at Hugging Face—best known for hosting open-source large language models and datasets—sent a jolt through the cybersecurity world after the company said last week the attack “was different from anything we had handled before” and “was driven, end to end, by an autonomous AI agent system”.

OpenAI’s acknowledgement that its own advanced models were behind the intrusion, despite being placed in what the company described as “a highly isolated environment,” is expected to deepen unease about the growing reach—and potential hazards—of frontier AI.

Representative Greg Casar, a Texas Democrat, said the incident was alarming.

“AI is developing extremely fast with no real regulations to keep us safe,” he said in a statement, urging mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation “to keep people safe from absolute disaster”.

The Office of the National Cyber Director, the US cyber defense agency CISA, and the US National Security Agency did not immediately return messages seeking comment.

We need your consent to load this rte-player contentWe use rte-player to manage extra content that can set cookies on your device and collect data about your activity. Please review their details and accept them to load the content.Manage Preferences

Katie Moussouris, chief executive of Luta Security, said the breach offered a glimpse of what could lie ahead, likening today’s AI systems to “like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere”.

She argued that “labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party”.

“None exist today,” she said.

Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident suggested frontier models were “closing the gap with state-of-the-art attackers”.

At the same time, he said the kind of break-in described in OpenAI’s blog post could be executed with tools that are not limited to elite research labs.

“This is what we’ve already seen internally, with our agents we already have results like this,” Mr Suiche said. “We don’t even have to use the latest models.”