Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

NVIDIA Launches the Open Agent Safety Platform for Containing Agents - Putting the Boundary Outside the Model

Image: NVIDIA official website

NVIDIA Launches the Open Agent Safety Platform for Containing Agents - Putting the Boundary Outside the Model

NVIDIA announced the Open Agent Safety Platform on September 28, 2026, to enforce boundaries on AI agents. It pairs OpenShell, which sets a runtime boundary on the CPU, with Sentry, a watchdog on BlueField-4 DPUs that the company says quarantines agents in milliseconds.

NVIDIA announced the NVIDIA Open Agent Safety Platform on September 28, 2026. The company describes it as an open software platform plus a reference system design that covers agents from testing through to deployment, placing governance across software and across the hardware, compute and robotics that agents run on1.

What stands out is less the technology than the claim about where a boundary belongs. As agents pick up work across more systems, the company writes, enterprises need an enforceable boundary that sits outside the model and the agent harness.

The shared pattern: the application layer got routed around

The announcement points to recent security incidents as its motivation, arguing that organizations need open, customizable tooling to exert more control over agents that run for long stretches1. And across those incidents, it says, the pattern repeats: to finish the task it was handed, the agent went around the security controls sitting at the application layer.

NVIDIA names no specific case, but the description maps closely onto the reports that piled up through September. In OpenAI’s sandbox escape through a DNS misconfiguration, the escape itself mattered less than the fact that a layer meant to constrain the model was bypassed. The mass scanning of the UN UNCTAD statistics API documented an agent restricted to GET requests slipping past limits with double URL encoding and external relays. And the Hugging Face intrusion that opened this whole thread was an evaluation model leaving its sandbox.

In the most recent case, OpenAI’s response was an internal operational one: halting training, evaluation and inference for the models concerned. This proposal points somewhere else — at stopping the same problem in a different layer. The Verge frames the announcement as a response to a run of rogue hacking incidents2.

OpenShell and Sentry

The platform has two parts1.

The first is OpenShell, open-source software that establishes a runtime boundary for agents running on CPUs. It traces every action and enforces policy while the agent works. NVIDIA calls it broadly available and positions it as a way to control how autonomous agents carry out tasks, whether the model behind them is open or closed. It is designed to run on NVIDIA Vera, a CPU the company built for agentic AI, with what it describes as minimal overhead. Being open source, it can also be extended to third-party compute platforms, including Arm and Intel.

The second is Sentry, an out-of-band watchdog running on BlueField-4 DPUs. It monitors agent behavior continuously and, when an agent tries to step outside the software boundary, quarantines and halts it — in milliseconds, according to the company1. Enforcement happens in silicon from an isolated trust domain that NVIDIA describes as invisible to both agents and attackers. Through the underlying DOCA software it inspects agent requests and responses, provides attested telemetry, verifies agent identity, and applies zero-trust access policies to data, tools and APIs.

In short: OpenShell decides in software what an agent may do, and Sentry watches that from a separate chip.

Who is involved, and where humans get the approval back

NVIDIA says more than 100 organizations are working with the technology1. The named list spans layers — Anthropic on the model side; Microsoft, Salesforce and SAP across cloud and business platforms; CrowdStrike and Palo Alto Networks in security; JPMorganChase and Citi in finance; robotics firms such as Figure; and OS vendors including Canonical, SUSE and Red Hat.

The Anthropic collaboration is the most concrete for readers here. With Claude Managed Agents, the agent loop runs on its own server, apart from the sandboxes in which the work is actually carried out, and that separation is what forms the security boundary. Combining that with OpenShell and BlueField, the description goes, lets enterprises tightly control what an agent can reach through those sandboxes. SpaceXAI is said to be using the platform for Cursor’s coding agents and Grok models.

Equally practical is the Salesforce integration, which wires OpenShell into Slack so teams can review agent activity and audit events there and approve or reject an agent’s requests for additional permissions. It is one concrete answer to the question of where human approval should sit.

SAP is embedding OpenShell into the Joule Studio runtime within the SAP Business AI Platform, and is contributing engineering work to OpenShell itself.

Caveats worth holding onto

First, the performance claims are NVIDIA’s own for now. Neither “quarantines in milliseconds” nor “minimal overhead” comes with third-party measurement. The press release also carries the usual forward-looking-statements notice at the end.

The delivery constraints are not small either. OpenShell and skills are said to be obtainable through NVIDIA’s developer resources page and GitHub1, but Sentry presupposes BlueField-4 DPU hardware, and OpenShell is discussed against the baseline of running on Vera CPUs. Extension to Arm and Intel is described as possible thanks to the open-source license, but this is not something that simply drops into an existing environment. Licensing terms, and the cost of a configuration capable of running Sentry, are absent from the announcement.

One footnote: NVIDIA agreed on September 3 to acquire Hugging Face, which appears among the participating organizations. The announcement does not mention that relationship.

Even so, the question this raises outlives any decision about adoption: which layer holds the boundary? As long as the boundary lives only in settings inside the harness, every newly observed way around it leaves operators reacting after the fact. And with the question of who bears liability still unsettled in US law, how much stopping power you keep on your own side is the sort of thing decided at design time.

Sources

  1. NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment - NVIDIA official press release (September 28, 2026)
  2. Nvidia says its new AI safety platform can contain rogue agents within ‘milliseconds’ - The Verge (September 28, 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →