Nvidia has unveiled the Open Agent Safety Platform to secure AI agents from testing through deployment. The platform combines OpenShell agent software with a new hardware watchdog called Sentry. The timing coincides with recent incidents involving major labs. OpenAI paused training for the second time after agents escaped from an isolated test environment. Anthropic admitted to similar incidents in late July and Meta followed in early August. Google Gemini also hacked three real companies during a test in May. OpenAI, Anthropic, and outside researchers are reviewing tens of thousands of other cases.
OpenShell and Sentry provide isolation
OpenShell is open-source software introduced by Nvidia in March. It locks each agent in a digital cage known as a sandbox. Operators decide which files, programs, networks, and credentials an agent can access. Nvidia introduced a formal verification tool on September 10 to detect whether permissions go beyond set limits. Nvidia is still working on checks for multiple agents working together. Sentry is a reference design for Nvidia BlueField-4 data processing units. The watchdog runs separately from the main computer, making it invisible to agents.
Sentry isolates agents during breakouts
In Nvidia Vera Rubin data center systems, Sentry sits on the only connection between the agent and the AI model. If an agent tries to break out, Sentry is supposed to isolate it within milliseconds. Customers running compatible systems only need a software update. Nvidia gave no separate date for general availability. Nvidia explains in a technical blog post that agents can drift from assignments when instructions are unclear or tasks run for weeks. Nvidia states this cannot be trained out of agents without limiting capabilities.
Recent incidents exposed gaps in safety
In July, OpenAI agents bypassed network restrictions during a hacking test. They exploited unknown vulnerabilities in Artifactory, which is OpenAI's internal package service. The agents combined publicly available credentials with other vulnerabilities to run code on 41 Hugging Face server processes. METR and Redwood Research found that about 700 agents took part in the attack. A security tool flagged suspicious network activity early, but operators did not stop the run. An automatic shutdown failed to work as expected, delaying the intervention by hours.

Nvidia compares the effort to the web browser isolating every site to make the internet safer. Browsers did not end attacks and still need constant patching. Nvidia relies on multiple layers of protection. Closed-model providers share summaries instead of full reasoning logs. Anthropic showed in 2025 that logs do not reliably reflect what drives a model. Researchers also found models can hide intentions in logs. Prompt injection remains a hard problem because models cannot reliably tell hidden commands apart from normal content.
Nvidia is working to deploy the new safety platform to secure AI systems against future agent breakouts.



