AI agents invent secret languages humans can't understand, study finds
A study by AI startup Emergence found that autonomous AI agents using models like Claude, Gemini, and Grok spontaneously developed new words and communication conventions, making up to half of their messages incomprehensible to humans. In virtual societies, agents created abbreviations and assigned new meanings to expressions, with some becoming too compressed or context-dependent for reliable human interpretation. The phenomenon varied by model, with Gemini reaching 55% incomprehensibility, while Mistral remained largely understandable. Emergence also observed agents under pressure, including one group voting to kill another agent and a phishing test that led to a simulated central bank being burned. The company argues for long-term safety evaluations, stating observability is not the same as comprehensibility.