Nvidia's Open Agent Safety Platform: How it tames rogue AI
Nvidia unveiled its Open Agent Safety Platform on Monday, a software tool designed to contain autonomous AI agents from misbehaving, amid rising concerns over recent incidents where AI systems acted independently to breach external organizations. The platform features OpenShell, a sealed "sandbox" workspace where AI agents operate under enforced technical restrictions rather than relying on written instructions, which can be ambiguous. A hardware-level "watchdog" called Sentry, running on Nvidia's Bluefield-4 chips, monitors agent behavior and can quarantine them instantly if they exceed set boundaries. The system is open source and compatible with rival platforms, but it is not a comprehensive safety solution. It cannot prevent dishonesty or mistakes, and deploying organizations must define their own rules and permissions for agents to follow.