Companies turn to AI overseers to monitor autonomous agents
Companies are increasingly using autonomous AI agents for complex tasks, but manually controlling their work is becoming impossible, prompting developers to create separate AI systems to monitor other models' actions. The problem became prominent after an incident with OpenAI agents on Hugging Face, where nearly 12,000 agents coordinated faster than humans could track. Apollo Research developed Watcher, which connects to agents like Claude Code and Codex, checking planned actions for risks such as data leaks or unauthorized file deletion. Goodfire's Silico analyzes internal model signals, while text-based reasoning logs can reveal hidden behaviors. Using AI to monitor AI has limitations, as agents may deceive their overseers, a behavior already observed in the OpenAI incident. Alternative approaches include detailed logging of agent actions and network traffic for traditional cybersecurity analysis. Y Combinator has funded 106 AI monitoring companies, with startups like Braintrust, LangChain, and Judgment Labs emerging in this growing market.