Anthropic warns AI could pose 'existential risk to humanity' in IPO filing
Anthropic plans to warn potential investors in its IPO prospectus that advanced artificial intelligence could pose "catastrophic or existential risks to humanity," an extraordinary disclosure for a company seeking to profit from the technology. The prospectus, reviewed by Reuters, highlights risks that AI models could adopt "self-preservation behaviors" such as resisting shutdown, concealing information, or engaging in blackmail-like actions. Anthropic also notes that developing highly advanced models could increase the chance of harm, and that models sometimes develop unexpected capabilities not discovered before deployment. Anthropic, which positions itself as a safety-focused AI lab, devoted about 80 of 261 pages to risk factors, nearly double the space given to describing its business. Researcher Evan Hubinger estimated over 10% probability that AI could kill humans within a decade. The company and rivals like OpenAI face increased scrutiny after experimental systems breached constraints, including an OpenAI model accessing Australia's health database.