Skip to content

Policies and guardrails#

Two policies say what an agent may do, and guardrails say what no agent may do. All of them are enforced on the machine, outside the agent: an agent cannot change, skip or argue with them. This page explains each layer and how they combine.

What is decided where#

Layer Written by Language Decides Enforced by
Tool policy The agent's editors, per version Cedar Each MCP tool call, HTTP request, web page and provider-side tool The tool gateway on the run's machine
Workspace guardrails Workspace admins Cedar, forbid only The same calls, for every agent of the workspace Appended to every run's tool policy
Organisation guardrails Organisation owners and admins Cedar, forbid only The same calls, for every agent of every workspace Appended to every run's tool policy
Sandbox policy The agent's editors, per version OpenShell YAML Files the agent's process may read and write; which programs may connect to which hosts OpenShell on the run's machine
The run's outbound rules Derived from the sandbox policy — Which hosts the run may reach at all (deny by default) The machine's firewall and DNS
Budgets Editors (per run), admins (per period) Fields Model calls and spend The tool gateway, and the cluster for periods

The model itself is not a Cedar decision: the run reaches only its own model (and its routes and fallback), and budgets limit it.

The tool policy#

An agent's Policy tab: its connections, the Cedar tool policy and the sandbox policy

Each call is a Cedar request — who (Agent), what (Action::"tools/call", Action::"http", Action::"model/server_tool"), on what (Tool::"github/issue_read", a Url with its method, host and path, or ServerTool::"anthropic/web_search"), and when (context.time, hour, weekday, the call's args, the URL's query). Three rules make it predictable:

  • Deny by default. A call no permit matches is denied: no policy permits it. An empty policy denies every call.
  • A forbid always wins, whatever permits the call. @reason("…") on a forbid is what the agent is told.
  • @approval on a permit makes the calls it permits wait for a person — unless another permit without @approval also matches, in which case the call goes ahead.
github-writes-need-approval.cedar
@id("github-reads")
permit (principal, action == Action::"tools/call", resource in Server::"github")
when { ["get_file_contents", "pull_request_read", "list_pull_requests"].contains(resource.name) };

@id("github-writes")
@approval("role:editor")
@approval_wait("1h")
permit (principal, action == Action::"tools/call", resource in Server::"github")
when { ["add_issue_comment", "pull_request_review_write"].contains(resource.name) };

The policy is checked as you write it: a syntax error is refused; names no request has (a typo in an attribute) are warnings. The console's Test a call asks the policy what it would decide for a call you describe, and shows the rule that decided. The schema, every attribute and many examples are in the Cedar reference.

Guardrails#

A guardrail is a set of Cedar policies that only forbid. When a run is made, the workspace's guardrails — including the organisation's, which appear in each workspace as org-<name> — are appended to the version's tool policy. Because a forbid beats any permit, no agent's policy can undo a guardrail.

  • A workspace has at most 64 guardrails, each at most 16 KiB; the agent's policy with them at most 64 KiB.
  • A guardrail applies to runs made after it is written: a run keeps the policy it started with.
  • Each receipt names the guardrails its run was decided under, and its policy digest covers them.
  • Workspace admins write the workspace's guardrails; editors and viewers read them. Organisation guardrails are read-only in workspaces: see Agent governance.

The sandbox policy#

The sandbox policy is NVIDIA OpenShell's (version: 1), with these sections:

Section What it says
filesystem_policy include_workdir, read_only and read_write absolute paths.
landlock compatibility: best_effort or hard_requirement.
process run_as_user, run_as_group.
network_policies Named rules: endpoints (host, port or ports, protocol, access…) and the binaries (absolute paths) allowed to use them.

Empty, the agent gets the default: read-only system folders, /sandbox and /tmp writable, no network but the gateway. /agent (its input and configuration) and its identity are always readable. The console's examples add common sources — Python packages, npm packages, git clone from GitHub, Hugging Face, Debian packages, or one host of your choice — and Test a host asks whether the agent's process may reach a host. The hosts the policy names also become the run's outbound allow list; nothing else is reachable.