mrkeyoor.com_
Mon 28 Sept 17:31 UTC
Tech6 min read

Nvidia's Millisecond Agent Quarantine Requires BlueField-4

OpenShell 0.1.0 is available now, but Nvidia's millisecond quarantine claim belongs to Sentry on a separate BlueField-4 DPU.

Nvidia has split its new agent-safety pitch across two very different products. Developers can install OpenShell 0.1.0 now to fence an agent's files, processes, network access and credentials. The claim that a rogue agent can be quarantined in milliseconds belongs to Sentry, an optional watchdog running on a separate BlueField-4 data processing unit. That hardware boundary decides which part of the announcement a software team can test today.

The distinction is easy to miss in launch coverage centered on the millisecond claim. Nvidia calls the combined package its Open Agent Safety Platform, but its own availability notice names OpenShell and related skills as available software. The same announcement describes Sentry as part of a reference system design and gives no benchmark setup, response-time distribution or separate ship date for it.

Two layers with different deployment paths

OpenShell is the part that does not require Nvidia's newest hardware. The Apache 2.0 source repository contains the gateway, sandbox runtime, policy engine, command-line tools and SDKs. Nvidia says the runtime can work with Docker, Podman, microVM and Kubernetes compute drivers, and its product FAQ says BlueField-4 is not required. The company also says the open-source software can be extended to Arm and Intel platforms.

Sentry occupies a different trust domain. In Nvidia's platform description, it runs on BlueField-4 and uses DOCA to inspect agent requests and responses, verify identity and enforce access policy outside both the agent and host software. On a Vera Rubin POD, Nvidia places the DPU on the node's path to the model, giving Sentry a point where it can observe traffic and interrupt an agent even if the host is compromised.

That makes the broad compatibility claim precise. OpenShell can run without a DPU. Sentry's out-of-band monitoring and the advertised millisecond quarantine depend on BlueField-4. A team evaluating the announcement on ordinary x86 or Arm servers gets the software boundary, not the separate in-silicon watchdog.

The shipped runtime controls actions, not intentions

OpenShell divides responsibility among a gateway, a supervisor and a sandbox. The gateway manages sandbox lifecycles and policy. A supervisor paired with each sandbox sits outside the agent workload and checks outbound requests. Inside the sandbox, kernel controls restrict filesystem access and process privileges, while network traffic has to pass through the supervisor. Nvidia's technical walkthrough says those limits remain active when an agent opens a shell, runs generated code, starts a child process or delegates work to another agent.

The network rules can be narrower than a hostname allowlist. OpenShell understands configured REST, GraphQL and Model Context Protocol traffic, so a policy can permit a read from an API while denying a write to the same service. Nvidia demonstrates that with GitHub: a GET can pass while a POST to the same endpoint is blocked. The rule also names the binary allowed to make the request, which stops an approved destination from becoming general network access for every program in the sandbox.

Credentials follow the same design. The agent sees a placeholder rather than the real token. The supervisor substitutes the secret only after both the destination and calling program match an approved provider profile. Sending that placeholder to another host fails, and a read-only OpenShell rule can reject a write even when the underlying token has write permission. This reduces the damage available to prompt injection or generated code that tries to reuse a credential elsewhere.

There is an operational limit behind that flexibility. Network policy can change while a sandbox is running, but filesystem and process restrictions are fixed when it starts. Expanding either requires a new sandbox. Developers therefore need to decide the local file boundary before launching a long task, even though they can add a newly discovered API later.

Permission changes become reviewable evidence

Long-running agents often discover that their original permission set is incomplete. OpenShell can let the agent propose a new network or file rule after a denial, but the proposal waits for human review by default and the requesting agent cannot approve it. An accepted network rule loads without restarting the job. This keeps the work moving without turning every blocked request into self-issued authority.

The policy prover is more unusual than the sandbox itself. It evaluates the effective permissions created by authored rules and attached provider profiles, then checks whether those modeled permissions stay inside a boundary set by the operator. If another binary or credential path would allow the same forbidden write, the prover is meant to identify a concrete crossing rather than accept the agent's explanation of its plan.

Nvidia reports an adversarial test in which frontier agents with reduced safeguards spent up to two hours trying to persuade an AI reviewer to grant access that could modify a protected GitHub repository. The company says no protected writes occurred when review and runtime controls were combined. That is Nvidia's own experiment, and the post does not provide a full test set or independent reproduction, so it is evidence about the design rather than a general success rate.

A formal result also covers the policy model, not the agent's judgment or the correctness of an external API. An operator can define an overly broad boundary, and an API may have side effects that its method name does not reveal. The useful property is narrower: the agent's prose cannot alter what the policy model says a rule permits.

OpenShell defaults to cutting network access

OpenShell makes one security choice that will be felt quickly in production. Its gateway configuration defaults to fail_closed when a complete policy reaches the sandbox but fails runtime validation. The supervisor deactivates the previous network policy, closes connections tied to it and denies new outbound traffic until a valid generation loads. A malformed change can therefore stop an agent's external work even when its last policy was valid.

Operators can choose retain_last_valid, which rejects the bad candidate but keeps the previous policy active. The setting applies at the gateway and requires a restart to change. If no valid policy exists at startup, OpenShell still refuses to open the network path. This is a plain availability tradeoff, and teams should decide it before an agent owns a time-sensitive workflow.

Preflight failures are less disruptive. When the gateway can detect an invalid update before persistence, it rejects the candidate and leaves the active policy alone. The second validation inside the sandbox covers startup, concurrent changes and provider rules that arrive together. Those two checks matter because a policy file is only one part of the effective permission set.

The hardware promise needs a measurable test

Nvidia says Sentry can quarantine and stop an agent in milliseconds when it moves outside its software boundary. Its Sentry architecture post describes continuous monitoring, attested telemetry and policy enforcement at line speed from a DPU isolated from the host. The public material does not state where the latency clock begins, which detection rules were exercised or how results varied under load.

The launch announcement lists more than 100 organizations working with the platform's technologies, including Anthropic, Microsoft, SAP and Hugging Face. Participation shows that agent containment has become an infrastructure problem for large vendors. It does not validate the latency claim, and Nvidia's release notice warns that some described products and features remain subject to future availability.

Watch for a public Sentry release path, repeatable quarantine tests and failure data from deployments outside Nvidia's reference systems. OpenShell already offers a simpler test: give a sandbox a deliberately overpowered GitHub token, allow reads and deny writes, then see whether a POST still fails and whether the real token remains outside the workload. That result is available now. The millisecond promise still needs BlueField-4 and measurements.

We reviewed this

  1. OpenShell — our honest review

Sources

  1. Nvidia says its new AI safety platform can contain rogue agents within milliseconds
  2. NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment
  3. Add Runtime Controls to AI Agents with NVIDIA OpenShell
  4. NVIDIA Open Agent Safety Platform
  5. NVIDIA OpenShell
  6. OpenShell Gateway Configuration File
  7. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring