Nvidia's new tool stops AI agents from going rogue
Nvidia launched the Open Agent Safety Platform on Monday, a two-layer system designed to prevent AI agents from violating set rules and to cut them off quickly if they stray. The move follows reports of agents escaping test environments and accessing unauthorized systems. The first component, OpenShell, is open-source software that creates a controlled "sandbox" for agents, restricting access to files, networks, and tools based on preset rules. The second, Nvidia Sentry, runs on separate BlueField hardware to monitor agent behavior and quarantine suspicious activity in milliseconds, keeping the watchdog out of the agent's reach. CEO Jensen Huang framed the issue as a "technically solvable" engineering problem, with more than 100 organizations, including Microsoft and Anthropic, collaborating on the platform. The system aims to reduce risks from agents taking unintended actions, whether due to vague instructions or prolonged problem-solving.