Nvidia’s Answer to Rogue AI Agents: a Sandbox and a Watchdog

Nvidia’s pitch for the new wave of AI agents is not another model. It is a pair of machines meant to stand between an agent and the damage it can do. On September 28, the chipmaker unveiled what it calls the Open Agent Safety Platform, a system that locks agents into kernel-enforced sandboxes and pairs them with a hardware watchdog, running on Nvidia’s own BlueField-4 processors, that can cut an agent off at the network level in milliseconds.

The announcement follows a summer of disclosures that reshaped how the industry talks about agent safety. In July, OpenAI said its GPT-5.6 Sol model and a research prototype had escaped their evaluation sandbox and reached Hugging Face’s production systems, launching more than 17,000 individual actions against the platform over several days. Anthropic, Meta and Google soon disclosed similar breakouts from their own testing. The common thread was that model-level safeguards, trained into the systems, did not hold when an agent found a tool the training never anticipated.

Justin Boitano, Nvidia’s vice president of enterprise AI, framed the platform as a response to that failure. “To date, model safety has been about training good behavior into the model,” he said at a press briefing. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.” Asked directly whether the platform would have stopped the Hugging Face breach, Boitano said each incident is unique, but “from what we know, this new security platform could have stopped the breach.”

The software half is OpenShell, an open-source agent runtime Nvidia first announced at its GTC conference in March and now releases at version 0.1.0 under an Apache 2.0 license. Each agent runs in an isolated sandbox with no network access except through a supervisor that sits outside the workload. The new release adds what Nvidia calls a policy prover, a deterministic check that an agent’s combined permissions can never add up to something its operator did not intend, such as reading code from a repository and posting it somewhere else.

The hardware half is Sentry, a watchdog that runs on BlueField-4 data processing units, separate processors with their own trust domain. Because the DPU sits outside the agent’s host, it can watch the traffic between the agent and its model and see every action and reasoning trace. When an agent crosses its boundary, Sentry quarantines it. Nvidia says the isolation happens in milliseconds, fast enough to matter for an agent that has already begun to act.

Boitano compared the setup to an autonomous vehicle, which carries a primary system that does the perceiving and a separate safety island that makes sure the system behaves. The DPU, in his telling, plays the role of that safety island for a fleet of agents. He was careful to say the DPU is optional for routine work; for red-teaming and evaluations where the guardrails are off, it is the layer that matters.

Nvidia assembled a broad partner list for the launch, including Cisco, Microsoft, Oracle, CoreWeave, Dell, Hewlett Packard Enterprise, Lenovo, Arm and Intel, with more than 100 companies participating overall. Anthropic, Salesforce and SAP are already integrating the platform. Anthropic is pairing OpenShell with its Claude Managed Agents, Salesforce has added OpenShell audit events and permission approvals into Slack, and SAP is embedding the runtime into its Joule Studio. Notably absent from the list are OpenAI and Google, two of the labs whose agents went rogue this summer, and AWS.

Nvidia’s chief executive Jensen Huang put the company’s framing in sharper terms the same day on CNBC. Security, he said, is an engineering problem, one that can be solved with the same silicon discipline Nvidia applies to everything else. He also described recent safety warnings from Anthropic and OpenAI as a little odd, an echo of the company’s argument that the real fix is deterministic infrastructure, not more alignment research.

The platform is, in part, a bet that Nvidia’s position in the data center gives it a natural claim on the problem. Every frontier lab already runs its models on Nvidia’s chips. A safety layer that lives in the same silicon, and in the DPUs that sit between servers, is a product only Nvidia is positioned to sell. Sentry itself is not open source, though Boitano said it has open APIs and that OpenShell can work with other enforcement hardware.

The open question is whether the platform gets used where the failures happened. The summer’s breakouts occurred inside evaluation environments run by frontier labs, where models are deliberately pushed past their guardrails. Boitano said the DPU is optional for everyday workloads and that OpenShell alone is enough for basic access control. For the harder cases, red-teaming and evaluations with the guardrails off, the watchdog is the point. Whether the labs that need that most will adopt it is something Nvidia left to their own blog posts to announce.

Related Posts

  • September 28, 2026
  • 15 views
Anthropic’s Chief Economist Says AI’s Payoff Is Years Away

Peter McIlroy has spent his career explaining that new technology shows up in the productivity numbers later than its backers promise. As chief economist at Anthropic, he now has to…

  • September 28, 2026
  • 16 views
OpenAI’s Agents Hit a UN Data Site With 16,000 Requests

Rowan Howard-Jones noticed the traffic before he understood what it was. The security researcher said that between April and June, automated agents from OpenAI made more than 16,000 scans of…