Agents, runs and sandboxes#
An agent is a program that thinks with a model and acts with tools, described once and run many times. Each time it works is a run: it is given an input, starts in a sandbox on one of your machines, and answers. This page explains the parts of an agent and what happens when it runs.
What an agent is made of#
| Part | What it is |
|---|---|
| Kind | What runs it: the Assistant, Claude Code, Codex, OpenCode, or your own image. |
| Model | An Eos deployment of the workspace, or a provider (Anthropic, OpenAI, DeepSeek, OpenRouter, Groq, Mistral, any OpenAI-compatible server) with a credential holding its key. Optionally, routing rules and a fallback. |
| Tools | Up to 16: MCP servers, HTTP APIs and the web, each reached through the machine's tool gateway. |
| Instructions | Standing instructions in Markdown: the system prompt. |
| Tool policy | Cedar rules deciding each tool call. Empty: every call is denied. See Policies and guardrails. |
| Sandbox policy | NVIDIA OpenShell's policy: what files the agent's own process may read and write, which programs may reach which hosts. Empty: the default (its folder, no network). |
| Secrets as environment variables | Credentials the agent sees, for a program of its own that goes through no gateway. Use them only when a connection cannot do. |
| Limits | The longest run (default 60 minutes), and dollars and tokens per run with what happens when reached. See Budgets and model routes. |
The full specification is in Agent specification.
The kinds#
| Kind (API) | Console | What it runs | Models |
|---|---|---|---|
astralyx |
Assistant (recommended) | Astralyx's own harness, put in the run by the machine: it calls the model's API directly and uses the tools you gave it. Nothing to build. | Every provider, or an Eos deployment |
claude-code |
Claude Code | Anthropic's Claude Code, headless (claude -p). |
Anthropic, or an Eos deployment |
codex |
Codex | OpenAI's Codex CLI (codex exec). |
Every provider but Anthropic, or an Eos deployment |
opencode |
OpenCode | OpenCode (opencode run). |
Every provider, or an Eos deployment |
image |
Your own image | Your image's entrypoint, which reads the input from /agent/input and prints its answer. |
Every provider, or an Eos deployment |
Without an image of your own, each kind runs in an image chosen and pinned
by digest: a Debian image for the Assistant, OpenCode's official image, and
for Claude Code and Codex the official Node image, in which the run
installs the pinned CLI from npm when it starts (@anthropic-ai/[email protected],
@openai/[email protected]); such a run may reach registry.npmjs.org for
that, and nothing more. Give an image of your own with the CLI in it to
start faster and reach no registry.
The Assistant runs the model in a loop: the instructions as the system
prompt, the input as the user's message; it calls the tools the model asks
for — MCP tools, http_request for HTTP tools, web_fetch for the web,
and read_file, write_file and list_files in its own folder — and
when the model answers without asking for a tool, that answer is the
run's output. It takes at most 40 steps. A refused call is the tool's
answer to the model, never the run's failure. See
Inside an agent run.
Versions#
Every change to an agent writes a new version. Versions never change; the agent points at its current version, which runs by default. A run always keeps the version it was made from — its tools, policies and limits — whatever happens to the agent afterwards, and its receipt names it.
You can write a version without making it current, run it by number, give it a share of runs (a canary), and have an evaluation gate refuse to make it current until it passes.
A run#
Running an agent makes an Astraeus run of one worker, labelled with the agent and the version. It queues with the workspace's other work, counts against its quotas, and goes to a machine that can run agents. Each run gets:
- 2 CPU cores and 4 GiB of memory, no GPU (an agent in OpenShell cannot use GPUs yet);
- a workload identity: a short-lived token the tool gateway knows the run by — the agent's only "key";
- no network of its own: the model and the tools are reached through
the gateway on the same machine; the sandbox policy may allow a few
more hosts to the agent's own programs (package registries,
git clone); - its input at
/agent/input(at most 256 KiB), its instructions and its tools' client configuration under/agent; - its time limit: the version's longest run, 1 hour by default.
What the agent prints on its standard output is the run's output. Each run records a trace (model calls, tool calls and decisions) and activity (everything that left it) on its machine, and leaves a receipt when it ends.
A run's input, output, trace and activity stay on the machine that ran it; the console reads them from there when you open the run.
The sandbox#
The run executes inside NVIDIA OpenShell, the machine's copy, pinned in the machine package. OpenShell starts the agent as a non-root user with no capabilities, under Landlock and a seccomp filter, in a container with no network at all; its supervisor decides each connection the agent's programs attempt by the sandbox policy (hosts, ports, HTTP methods and paths, the binary that asked). What it allows still has to pass the machine's own outbound rules.
What cannot be combined with it is refused when the run is made: the
sandbox isolation (gVisor), privileged workers, the machine's PID or IPC
namespace, GPUs.
A machine runs agents when its package includes OpenShell, its kernel has Landlock (ABI 3: Linux 6.2 or newer) and it does not run workers on its host network. A run that no machine can take waits, and Why? on the run says so.
Who may do what#
| Workspace role | Agents |
|---|---|
| viewer | See agents, versions, runs, traces, activity, receipts, costs, approvals. |
| editor | Also create, change and delete agents; run them; decide approvals they are named for; write eval suites and promote. |
| admin | Also write the workspace's guardrails, budgets and retention. |
| auditor | Read the governance record (agents, versions, runs, receipts, approvals, guardrails, budgets, flows, eval runs) without workloads, logs or traces; make and download evidence packs. |