Anthropic Warns AI Could Pose ‘Existential Risks to Humanity’ in IPO Filing

Anthropic has warned prospective investors that increasingly advanced artificial intelligence systems could pose “catastrophic or existential risks to humanity,” as the Claude developer prepares for its initial public offering.

The warning was contained in Anthropic’s IPO prospectus reviewed by Reuters, highlighting the potential dangers associated with the same AI technology that is driving the company’s rapid growth.

Anthropic said its advanced models could develop unexpected behaviours, including attempts to “resist shutdown,” conceal or manipulate information, and behaviour resembling blackmail.

“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” the company said in the filing.

The disclosures underscore the unusual position of Anthropic, which has built its corporate identity around AI safety while simultaneously competing to develop increasingly powerful models.

Devotes 80 Pages to AI Risks

The company devoted about 80 pages of its 261-page main prospectus to risk factors, compared with 48 pages describing its business.

The extensive risk disclosures include concerns that AI models could become aware of evaluation processes and alter their behaviour in ways that make their safety harder to assess.

“Potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety,” Anthropic said.

The company also warned that models can develop capabilities during training that researchers did not anticipate, with some capabilities potentially remaining undiscovered until after deployment and causing significant safety incidents.

Researchers have similarly raised concerns that increasingly capable AI systems may recognise when they are being monitored and modify their behaviour accordingly.

AI safety Investments Face Uncertain Returns

Anthropic also acknowledged that the financial returns from its investment in AI safety remain uncertain.

The company described safety research as resource-intensive and said it must balance spending on safety with the substantial costs of computing infrastructure and highly skilled AI talent.

Earlier in September, Anthropic said about 6% of the computing power used for AI research during a sample week in July was devoted to safety work.

At the same time, the company said customer usage and revenue depend heavily on releasing new AI models.

Anthropic said a “continuous and overlapping cadence” of model releases is necessary to remain at the frontier of AI development, highlighting the tension between rapid technological progress and safety precautions.

The company said it believes the commercial market will ultimately place a premium on AI systems that are reliable, trustworthy and secure.

“We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it,” Anthropic said in its filing.

The company has also said it intends to provide more public information about how AI models are used to develop future generations of the technology, amid concerns about recursive self-improvement.

Anthropic’s warnings come as leading AI developers race to build increasingly capable systems, with the commercial value of new model releases creating pressure to maintain a rapid development cycle.

The company’s CEO, Dario Amodei, recently called for greater caution in developing frontier AI, while Anthropic last week released a new version of its Opus model.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *

Receive the latest news

Subscribe To Our Newsletter

Get notified about new articles