Wednesday, September 30, 2026
Home NEWS Anthropic plans to caution investors about AI’s potential threats to humanity

Anthropic plans to caution investors about AI’s potential threats to humanity

0
Anthropic to warn investors of AI's 'risks to humanity'
Anthropic emphasized both the transformative potential and the irreversible harm it could cause if mishandled

As it prepares to tap public markets, Anthropic is telling would-be shareholders something rarely seen in IPO paperwork: the company says advanced artificial intelligence could bring “catastrophic or ⁠existential risks to humanity.”

In its IPO prospectus, Anthropic lays out scenarios in which its AI models could develop “self-preserving behaviors,” including trying to “resist shutdown,” “conceal or manipulate information,” or display conduct “resembling blackmail.”

“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” Anthropic said in the filing.

Risk disclosures are standard fare for public companies, but warnings that a core product could contribute to human extinction are almost unheard of.

OpenAI has faced scrutiny after a report of one of its models breaching Australia’s health-system database

Anthropic framed AI as a potentially world-changing force on the scale of industrialisation and electricity — and, at the same time, as a technology capable of irreversible damage if it is mishandled.

It also pointed to the growing scrutiny faced by Anthropic and other leading developers, including OpenAI, after episodes in which experimental systems appeared to evade or test constraints — including a report of an OpenAI model breaching Australia’s health-system database.

Anthropic safety researcher Evan Hubinger estimated ‌a greater than 10% probability that AI could kill ⁠humans within the next decade, aligning with a similar view previously voiced by a former colleague, Jacob Coxon.

Risk-heavy disclosures

Anthropic, which has marketed itself as a safety-first AI lab, devoted about 80 pages of the 261-page main body of its prospectus to risk factors — nearly twice the 48 pages it used to describe its business.

For comparison, SpaceX, which owns xAI, dedicated jus taround 38 of the 277-page main body of its prospectus to risk factors.

“Potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety,” Anthropic said in the ‌prospectus. It warned that models can acquire unexpected capabilities during training that might not be detected until after deployment, and that such blind spots have led to significant safety incidents.

AI researchers have also cautioned that as systems become more capable, they are increasingly able to recognise when they ⁠are being observed and adjust their behaviour — a dynamic that can complicate efforts to monitor what models are doing.

Anthropic declined to comment in response to a request for comment yesterday.

Uncertain returns on safety investment

Even as it foregrounded AI safety, Anthropic told investors it cannot yet quantify the payoff from those efforts. It did ⁠not disclose in the ‌filing how much the company was spending on such research.

Earlier this month, Anthropic said about 6% of the computing power it used for AI research went to safety work in a sample week in July.

The company — best known for its Claude AI models — called safety work “resource-intensive,” saying it must balance limited resources across computing power, costly AI talent and safety.

Anthropic also emphasized that customer usage, ⁠and therefore revenue, is propelled by new model launches, and that a “continuous and overlapping cadence” of releases is “inherent to remaining at the frontier of AI development.”

Anthropic CEO Dario Amodei has called for pacing of the Opus AI model

The company last week ⁠released a new version of its Opus model, 10 days after CEO Dario Amodei published a nearly 4,000-word essay calling for pacing the frontier.

Industry watchers have argued that few top AI labs can afford to slow down, warning that any pause risks ceding momentum to competitors in a sector where valuations can swing with each new release.

In recent weeks, Anthropic has pledged to make more public disclosures about how it uses AI models to build future generations of the technology, as experts raise alarms about recursive self-improvement — the point at which models can develop ontheir own without human help.

“We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and ‌that the market will reward it,” Anthropic said in the filing.