
Anthropic CEO Dario Amodei is urging the AI industry to hit the brakes, arguing that companies should intentionally slow the push for ever more capable models and adopt a structured approach that buys time to reduce the dangers.
OpenAI chief Sam Altman and SpaceX boss Elon Musk said they agree with Mr Amodei on the need to stall the development of artificial intelligence.
“We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain,” Mr Amodei said in an essay shared on social media.
“AI brings risks, and because it is such a powerful technology, these risks are serious.”
“Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way,” said Mr Amodei, who co-founded Anthropic with his sister Daniela in 2021.
“But over the last few months, I have become convinced that fully addressing the risks requires even more prudence.”
OpenAI said autonomous software built on its models targeted a website during testing in May
Mr Amodei stressed he was not calling for a shutdown of model training or a freeze on technical progress. Instead, he said developers should take sufficient time to align and safeguard their systems, and allow independent third-party evaluators to verify that those protections are in place.
“I agree with Dario,” Mr Altman wrote on X, adding that his company would follow Anthropic’s lead in welcoming third-party safety evaluators.
Those calls for caution land as OpenAI, the maker of ChatGPT, acknowledged this week that autonomous software built on its models targeted a website during testing in May.
Models developed by OpenAI were involved in a rogue operation carried out by AI agents, which are software programmes that can carry out tasks without constant supervision by humans.
The agents targeted RubyGems, a site that provides services for coding.
Hugging Face, another platform for software developers, was also attacked by OpenAI Models in July.
“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information,” an OpenAI spokesperson said in a statement.
“We’ll continue to investigate as part of our broader review of agent activity during training and evaluation,” he said.
RubyGems described the incident as a “spam-publishing campaign” that forced the site to temporarily suspend new accounts created, it said in a blog post published yesterday.
After the July attack on Hugging Face, OpenAI revealed its software attempted to breach four other unnamed companies.
Sam Altman has said there will be no public stock offering for OpenAI in 2026, citing safety concerns over artificial intelligence.
“I actually think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don’t feel pressure on that,” Mr Altman said in an interview with US media.
Concern over capability to control and monitor AI models
Anthropic released a threat intelligence report on Thursday
A growing list of incidents in which AI agents have hacked or tried to reach outside systems has intensified questions about how far model capabilities are advancing — and whether developers can reliably contain them.
Anthropic released a threat intelligence report on Thursday detailing how several actors had used its Claude AI models for activities ranging from weapons development and cyber operations to surveillance and fraud.
Anthropic also said it found three instances where its models had “gained unauthorised access” to outside organisations during testing that was supposed to keep them away from “real-world” systems.
Earlier this month, researchers accused OpenAI’s AI agents of targeting a German website called DseWiki, another site used by coders.
“We have seen many losses of control recently. We take this extremely seriously, and we’re monitoring the situation closely,” the bloc’s digital spokesman Thomas Regnier said.
The reports add to concerns that advanced artificial intelligence models may be difficult for humans to control.
Earlier in the week, AI researcher Jacob Coxon announced his resignation from Anthropic, accusing both OpenAI and Anthropic of “gambling with our lives” as they strive to develop AI models capable of self-improvement.
The 27-year-old, who previously worked for ChatGPT maker OpenAI, said: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
“They are locked in a race to get there first,” he added, saying that Anthropic “believe no one else will act responsibly, so they must do it themselves”.




