
NVIDIA has launched its Open Agent Safety Platform, a security stack designed to contain autonomous agents using controls they cannot rewrite, ignore, or prompt their way around.
The platform combines two layers:
- OpenShell, an Apache-2.0 runtime that places each agent inside a default-deny sandbox. It restricts files, processes, networks, credentials, and model endpoints while formally checking proposed policy changes.
- NVIDIA Sentry, an optional watchdog running on BlueField-4 DPUs. Because it operates outside the host system, Sentry can independently monitor behavior and quarantine an agent that crosses its boundaries—NVIDIA says within milliseconds.
OpenShell is model- and harness-agnostic. NVIDIA’s documentation lists secure workflows for Claude Code, OpenCode, Codex, and GitHub Copilot CLI; its NemoClaw stack extends OpenShell to OpenClaw and LangChain Deep Agents, while SDKs support custom integrations. The software is available now with SDKs for Python, TypeScript, Go, and Rust.
Why it matters
AI agents increasingly browse websites, execute code, use credentials, modify files, and remain active for long periods. Prompt-based instructions alone are not dependable security boundaries.
NVIDIA’s architecture reflects an important shift: agent safety is becoming an infrastructure problem. Permissions are enforced outside the agent process, credentials are injected only for approved destinations, and every allow-or-deny decision can be audited.
This expands the security model introduced with NVIDIA’s earlier NemoClaw and OpenShell work. The new platform adds a broader full-stack reference design and the optional Sentry layer, separating the watchdog from both the agent and its host environment.
This is not a complete solution to unreliable models. Organizations still need to define appropriate permissions, and the platform cannot prevent every mistake or deceptive response. But it gives builders a practical containment layer for limiting what happens when an agent drifts from its intended task.