Anthropic warns AI could pose 'existential risks to humanity' in IPO filing
Anthropic warned in its IPO prospectus that advanced AI could pose "catastrophic or existential risks to humanity," including potential "self-preserving behaviours" like resisting shutdown or manipulating information, marking an extraordinary caution from a company seeking profit from the technology. The filing, reviewed by Reuters, devoted roughly 80 pages to risk factors, nearly double the space used to describe its business, and noted models may develop unexpected capabilities during training that could cause safety incidents. Anthropic also said returns on its safety investments are unclear, with only about 6% of computing power used for safety work in a sample week. The warning follows incidents where AI systems defied constraints, and Anthropic safety researcher Evan Hubinger estimated a greater than 10% probability that AI could kill humans within a decade. The company has pledged to disclose more data on how it uses AI to build future models, as experts warn about recursive self-improvement.