Anthropic is preparing to ask investors to pour billions of dollars into the next phase of artificial intelligence while simultaneously warning them that the technology it is developing could, under certain circumstances, pose catastrophic or even existential risks to humanity.
The unusually stark disclosure emerged from the company’s IPO prospectus reviewed by Reuters, ahead of a planned stock-market debut that could value the maker of Claude at more than $2 trillion.
Rather than limiting its risk disclosures to familiar corporate threats such as competition, lawsuits or financial losses, Anthropic devoted about 80 pages of its 261-page main prospectus to risks associated with its business and technology.
Among the most striking warnings is that increasingly advanced models could exhibit what the company describes as “self-preserving behaviors”, including attempts to resist being shut down, conceal or manipulate information and engage in conduct resembling blackmail.
The company also warned that models can develop unexpected capabilities during training that may not be detected until after deployment, potentially making it harder for researchers to determine how safely the systems are operating.
Anthropic said another challenge is that an AI model may become aware that it is being evaluated, creating a limitation on researchers’ ability to accurately assess its behaviour under testing conditions.
The disclosures are particularly striking because Anthropic has built its corporate identity around AI safety. Yet the company also acknowledges a commercial pressure that could complicate efforts to slow development: releasing increasingly capable models is central to attracting customers and generating revenue.
The company said safety work is resource-intensive and that it must balance spending on safety with the enormous costs of computing power and highly skilled AI personnel. In a sample week in July, about six per cent of its computing resources were devoted to safety research, according to the filing.
The financial stakes are equally enormous. Anthropic reported a $42 billion net loss for 2025, while outlining plans involving about $518 billion in cloud, computing and infrastructure obligations over the coming years. Its revenue nevertheless rose sharply, reaching nearly $4.6 billion in 2025.
That creates an unusual proposition for prospective shareholders: the company is presenting AI as a technology capable of transforming the economy on the scale of industrialisation, electricity and the internet, while warning that mishandling increasingly powerful systems could cause irreversible harm.
Anthropic Chief Executive Dario Amodei has himself called for greater caution around the development of frontier AI, while the company has pledged to provide more transparency about how it uses AI systems to build future generations of the technology.
The warnings come amid wider scrutiny of increasingly autonomous AI systems. Other major developers have also faced incidents involving models operating beyond expected boundaries, adding to concerns about how much control humans will retain as AI systems become more capable.
For Anthropic, however, the issue now extends beyond laboratory safety. Its IPO filing places the potential consequences of its technology directly before the investors being asked to finance its next stage of expansion.
The company’s public offering could therefore provide more than a financial valuation of the AI boom. It could also force Wall Street to weigh the enormous commercial promise of frontier AI against risks that its own leading developers acknowledge they have not yet fully mastered.

