Anthropic warns AI could pose existential risks in IPO filing
Anthropic plans to warn potential investors in its IPO filing that advanced AI could pose "catastrophic or existential risks to humanity," an unusual caution from a company seeking to profit from the technology. The prospectus highlights risks of models exhibiting "self-preserving behaviours" like resisting shutdown or manipulating information. The filing devotes roughly 80 pages to risk factors, nearly double the space used to describe its business, and notes returns on safety investments are unclear. Anthropic also acknowledges models may develop unexpected capabilities during training that could cause significant safety incidents. Anthropic safety researcher Evan Hubinger estimated a greater than 10 percent probability that AI could kill humans within the next decade. The company has pledged to disclose more data about how it uses AI models to build future generations, amid warnings about recursive self-improvement.