Guides/Playbook

NVIDIA launches Open Agent Safety Platform: Verify the policy

6 min read
On this page
An allowed permission footprint stays inside a review boundary while a proposed extension crosses it

OpenShell policy boundaries

General Analysis

NVIDIA launched Open Agent Safety Platform on September 28, 2026, combining OpenShell runtime controls with the Sentry hardware reference design. For a security team considering deployment, the immediate question is what authority an agent will actually receive. A sandbox can faithfully enforce a policy that grants too much access.

Start the review with the effective policy and the people or services allowed to change it. OpenShell provides a proposal risk check and a separate check against an operator-defined permission boundary. Passing the first does not establish the second. That distinction matters when an agent asks for more access to finish a task.

This guide interprets NVIDIA's documentation; General Analysis has not tested the deployment or reproduced a sandbox escape. It extends our coding-agent rollout baseline with an OpenShell-specific approval procedure.

Separate the runtime from the hardware layer#

OpenShell runs the agent workload inside a sandbox. NVIDIA's runtime walkthrough describes a supervisor outside that workload checking outbound requests, with a gateway managing sandbox lifecycles and policies. Provider credentials can remain outside the workload and be substituted into authorized requests. The receiving service still applies the credential's own permissions.

Sentry is an additional layer in the reference architecture, using BlueField-4 and DOCA for independent monitoring and enforcement. Installing OpenShell on a workstation does not establish that this hardware layer exists. Record the components your deployment actually uses before assigning them control credit.

The launch walkthrough describes OpenShell 0.1.0; the current documentation reviewed here is labeled 0.1.2. Match these checks to your installed release. The settings below describe the current rollout procedure; some predate the platform launch.

See how your AI systems hold up under real attacks

General Analysis maps AI applications and agents, red teams prompts, retrieval, tools, MCP servers, browser actions, permissions, and business workflows, then turns findings into evidence your team can reproduce and retest.

Review what an allowed API route permits#

Consider a hypothetical agent that summarizes issues in one private repository. Giving it access to api.github.com answers only part of the authorization question. The review also needs the allowed executable, methods, paths, and credential scope.

NVIDIA's security guidance distinguishes destination access from inspected request rules. An endpoint without a configured inspection protocol does not restrict HTTP methods and paths. For inspected requests, enforcement: audit logs violations while forwarding traffic; enforcement: enforce blocks requests outside the rules.

For this hypothetical task, require the intended repository paths and read operations, then check every other route to the same service. A narrow rule offers little reassurance if another permitted route grants the write authority you meant to withhold. Keep the upstream credential narrow too: a proxy rule and a service permission have different owners and failure modes.

A clean proposal check is not your access boundary#

The Policy Advisor lets a sandboxed agent propose network rules. It is off by default; when enabled, proposals wait for review by default. Optional automatic approval accepts proposals only when the risk check and destination checks find nothing to flag. OpenShell also creates drafts from blocked connections even when the agent-facing advisor is disabled, and automatic approval applies to those drafts too.

Review the approval mode even when the advisor is off. Preserve the proposed rule, the decision, and the policy it changed. An agent's explanation can help a reviewer understand a request, but it cannot define the organization's maximum permitted access.

The prover documentation explicitly separates proposal risk checks from boundary checks. A proposal can pass its risk check and still exceed the boundary an operator would choose.

For the repository-summary task, have the task owner define the boundary before the agent requests an exception. Copying an overly permissive policy into the boundary file would simply preserve that access. If the agent asks to write an issue, the approver must decide whether that changes the assignment.

Inspect the composed policy before checking it#

Use a trusted operator environment with the installed OpenShell CLI and openshell-prover, plus a separately reviewed boundary.yaml. Substitute the actual sandbox name for review-agent. These read-only inspection commands are adapted from NVIDIA's documentation; they have not been run against a General Analysis deployment.

Code source: OpenShell policy inspection and boundary checks.

Shell
openshell policy get review-agent --base openshell policy get review-agent --full openshell sandbox get review-agent --policy-only > effective-policy.yaml openshell-prover check effective-policy.yaml --boundary boundary.yaml --output json openshell policy list review-agent

The base policy is the operator's editable policy. The effective policy includes provider-contributed rules. Review that composed result; checking only the original YAML can omit authority added by attached providers. Keep the exported file with the reviewed boundary and the checker result.

Only within_boundary passes the boundary check. Treat exceeds_boundary, error, unsupported, and inconclusive as unresolved or failed checks. Current coverage includes filesystem, process, Landlock, network connections, and REST rules. Protocols such as GraphQL, WebSocket, and MCP can return unsupported; runtime support for a protocol does not imply prover coverage.

A passing result is relative to the supplied boundary and modeled features. It does not establish that the task needs those permissions or that the running sandbox enforces them.

Close the gap between approval and activation#

NVIDIA's policy-management reference distinguishes gateway acceptance from sandbox activation. Record the intended revision as Loaded, then inspect its effective policy. Even a successful --wait can accompany an unchanged policy or a revision superseded before loading.

Review itemEvidence to retain
Intended authorityTask owner's boundary and permitted service operations
Provider-added accessEffective policy and attached-provider inventory
Permission expansionProposed rule, approver, decision, and resulting revision
Boundary verificationResult, modeled coverage, and any unresolved protocol
Runtime activationIntended loaded revision and matching effective policy
EnforcementHarmless allowed and denied requests against a controlled test endpoint; corresponding OpenShell records

Use a disposable environment and a service you control for the final check. Confirm the denial comes from OpenShell rather than a missing client or the destination's authentication error. A service returning 403 does not identify which control stopped the request.

Network changes can load while a sandbox runs and close connections opened under the old rules. Filesystem and process changes require recreating the sandbox; saving a revision is insufficient. Preserve needed state before replacement.

Approve sensitive work when the evidence describes the same task, boundary, provider set, and active revision. If one changes, reopen that part of the review. For an existing fleet responding to a disclosed flaw, keep this policy review separate from verifying an installed security update.

Browse all