Run specification#
A run is the object you send to POST …/jobs under a workspace's cluster path (see REST API); the console's New run form and astra build it for you. This page lists every field Astraeus accepts, with its type, default and limits, and what it does. Use it when you write a run by hand, or when Astraeus refuses one and names a field.
In the API a run is a job and a worker is a task. The field names keep those words.
Shape of a run#
{
"metadata": { "name": "train-resnet", "labels": { "team": "vision" } },
"spec": {
"task_template": {
"image": "nvcr.io/nvidia/pytorch:24.08-py3",
"command": "python",
"args": ["train.py", "--epochs", "90"],
"requested_resources": {
"cpu_cores": 8,
"memory_bytes": 68719476736,
"gpu_requests": { "count": 1 }
}
}
}
}
metadatanames the run.specsays what it runs (task_template, or severaltask_groups) and how its workers start, fail and end.- The API reads JSON. If you keep specs in YAML, convert them as you send them (
yq -o=json run.yaml). See Submit a run. - Unknown fields are refused, not ignored. A typo such as
spec.replicasorgpu_requests.modelreturns400 UNKNOWN_FIELDand names the field. The check coversspecandspec.task_template.requested_resources. - Sizes are integers in bytes; durations are integers in seconds.
A complete example#
The run below uses most options. Each numbered marker explains its line. Convert it to JSON before you send it.
metadata:
name: llama-finetune # (1)!
labels: { team: nlp }
spec:
start: Gang # (2)!
on_failure: RestartJob
lifetime: Batch
priority: 10 # (3)!
task_template:
image: registry.example.com/nlp/train:2026-09
pull_policy: IfNotPresent
registry_secret_ref: { external_secret_name: registry-login } # (4)!
command: torchrun
args: ["--nproc-per-node=8", "train.py", "--config", "configs/70b.yaml"]
env: { NCCL_DEBUG: INFO, HF_HOME: /data/hf }
restart_policy: OnFailure
time_limit_seconds: 86400 # (5)!
requested_resources:
gpu_requests:
count: 16 # (6)!
models: [H100, H200]
min_memory_gb: 80
healthy_only: true
machines: 2 # (7)!
per_gpu: { cpu_cores: 12, memory_bytes: 128849018880 }
network: { interconnect: infiniband, min_gbps: 200 }
topology: { keep_within: rack, patience_seconds: 600 }
node_selection:
mode: Any
match_labels: { pool: h100 }
datavolume_refs:
- { name: datasets, mount_path: /data, mode: ReadOnly }
- { name: checkpoints, mount_path: /ckpt, mode: ReadWrite }
secret_refs:
- external_secret_name: hf-token
env_vars: { token: HF_TOKEN }
security_context:
ulimits: [{ name: stack, soft: 67108864, hard: 67108864 }]
- Letters, digits,
-and_; at most 57 characters. The workers are namedllama-finetune-0andllama-finetune-1. - Every worker is placed and started together. A failure restarts all of them.
- At most the workspace's maximum priority, which is
0unless an admin raised it. - A credential whose value is JSON:
{"username": …, "password": …, "server": …}. - The workers are stopped and the run fails after 24 hours of running. A time limit also lets the run backfill.
- The run's total. With
machines: 2, each of the two machines gets 8. - Leave it out to let the scheduler pick the fewest machines on the best network.
Minimal examples#
{"metadata": {"name": "pretrain"},
"spec": {"start": "Gang", "task_template": {
"image": "nvcr.io/nvidia/pytorch:24.08-py3", "command": "bash",
"args": ["-c", "torchrun --nnodes=$WORLD_SIZE --node-rank=$RANK --nproc-per-node=$ASTRAEUS_GPU_COUNT --master-addr=$MASTER_ADDR --master-port=$MASTER_PORT train.py"],
"requested_resources": {"gpu_requests": {"count": 16},
"per_gpu": {"cpu_cores": 12, "memory_bytes": 128849018880}}}}}
metadata#
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | required | The run's name, unique in its workspace on that cluster. Letters, digits, - and _; no - at either end; no .. At most 57 characters; with task_groups, at most 57 minus the length of the longest group name minus 1, so that every <run>-<group>-<rank> stays a DNS label. Lowercase is recommended: the console only accepts [a-z0-9-]. |
labels |
map of string to string | {} |
Your labels. Copied onto every worker. Keys under astraeus.io/ are reserved and refused. |
The cluster adds astraeus.io/created-by (who submitted the run). Workers also carry job.astraeus.io/name, job.astraeus.io/replica-index, job.astraeus.io/group and job.astraeus.io/group-replica-index.
spec#
| Field | Type | Default | Description |
|---|---|---|---|
task_template |
worker template | — | What every worker runs. Required unless task_groups is set; must be left out when it is. |
task_groups |
list of worker groups | [] |
Several kinds of worker in one run: a launcher and its workers, an engine and its clients. |
leader_group |
string | the first group | With task_groups: the group whose rank-0 worker leads (its machine is MASTER_ADDR). |
complete_with |
string | — | With task_groups: the group whose completion completes the run. The other groups are then stopped, and one of them failing for good fails the run. Not with lifetime: Service. |
start |
Gang, MinAvailable or Independent |
see defaults | How workers are admitted and started. Gang: every worker placed and started together, or none. MinAvailable: starts once min_available workers can be placed together; the rest join as room appears. Independent: each worker when it fits. |
min_available |
integer ≥ 1 | — | Required with start: MinAvailable; refused otherwise. |
on_failure |
RestartJob, RestartTask or FailJob |
see defaults | What a worker's failure does. RestartJob: restart every worker. RestartTask: restart the one that failed; in a gang, the leader's failure still restarts all. FailJob: no restart; the run fails and its other workers are stopped. |
lifetime |
Batch or Service |
Batch |
Batch runs to completion. Service runs until you delete it: every worker is restarted whenever it ends (restart_policy becomes Always), and the run never completes. |
policy |
Gang, GangIndependent or Independent |
Independent |
The older form of start and on_failure, still accepted. When you set either of those, the cluster rewrites policy to match. |
priority |
integer, −1000 to 1000 | 0 |
Place in the queue: higher first, and may preempt lower. Capped by the workspace's maximum priority. See Priorities. |
array |
object | — | Run the template size times as independent copies. |
model |
object | — | A model of the workspace mounted in the workers. |
shared_volumes |
list | [] |
Stored and validated, but not provisioned in this version: no drive is created. Use a drive and datavolume_refs. |
output_volumes |
list | [] |
As shared_volumes, with retain defaulting to true. Not provisioned in this version. |
idle_auto_delete |
object | — | Stored and validated (enabled, idle_timeout, grace_period_at_start, warn_before_kill, thresholds, bypass_during_hours as "HH:MM-HH:MM <IANA zone>"), but nothing acts on it in this version. |
retry_on_heartbeat_expired |
boolean | false |
Accepted for compatibility; not read. A worker lost with its machine is restarted according to its restart_policy. |
max_heartbeat_retries |
integer, 0 to 1000 | 0 |
Accepted for compatibility; not read. |
Defaults for start and on_failure#
| You set | start becomes |
on_failure becomes |
|---|---|---|
neither, and no policy |
Independent |
RestartTask |
start: Gang only |
Gang |
RestartJob |
start: MinAvailable or Independent only |
as set | RestartTask |
policy: Gang |
Gang |
RestartJob |
policy: GangIndependent |
Gang |
RestartTask |
The API and the tools differ
The API's default is Independent. The console sets Gang for you when the run may take several machines (more than one GPU and no array, or more than one machine). astra astraeus run and astra slurm sbatch do not set start, so their runs start Independent. In specs you send to the API, set start yourself.
Rules checked when a run is created#
lifetime: Servicecannot have anarrayoron_failure: FailJob.start: Gangcannot havetask_groupswithdepends_on: a gang starts every worker together.- An
arrayneedsstart: Independent(or nostart) and cannot havetask_groups. - Group names are unique.
depends_on,leader_groupandcomplete_withname groups that exist. A group cannot depend on itself. - An enabled external access name is unique across the run's groups.
array#
| Field | Type | Default | Description |
|---|---|---|---|
size |
integer, 1 to 10 000 | required | How many workers. Each has the template's whole ask. Index i gets ASTRAEUS_ARRAY_TASK_ID=i (from 0) and ASTRAEUS_ARRAY_SIZE. |
max_parallel |
integer, 0 to size |
0 |
At most this many hold machines at once. 0: no cap. |
See Arrays.
model#
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | required | A model of the workspace. |
mount_path |
absolute path | /model |
Where its weights appear, read-only. ASTRAEUS_MODEL_PATH names it, and ASTRAEUS_MODEL_FILE names a single GGUF file. |
engine |
boolean | false |
Also serve the model beside the run: an engine worker group serves it, your workers become the main group with OPENAI_BASE_URL and OPENAI_MODEL set, and the run completes when main does. Not with lifetime: Service or an array. |
A cluster that does not serve models refuses a run with model (400 MODELS_NOT_SERVED). See Eos.
task_groups: worker groups#
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | required | The group's name. Its workers are named <run>-<name>-<i>. |
task_template |
worker template | required | What this group's workers run. |
depends_on |
list of {group_name, condition} |
[] |
Groups that must be up before this one is placed. condition: Running (the default) or Healthy (running and passing its health_check). Not with start: Gang. |
Ranks are global across groups, in the order the groups are listed. With a group worker of two workers and a group ps of one, the workers are run-worker-0 (rank 0), run-worker-1 (rank 1) and run-ps-0 (rank 2).
task_template: the worker template#
| Field | Type | Default | Description |
|---|---|---|---|
image |
string | required | The container image. busybox means docker.io/library/busybox:latest. |
pull_policy |
IfNotPresent, Always or None |
IfNotPresent |
When the machine pulls the image. None: never; the worker fails if the image is not already on the machine. |
registry_secret_ref |
{external_secret_name} |
— | A credential whose value is JSON {"username", "password", "server"}, used to pull from a private registry. |
command |
string | the image's | One executable. It replaces the image's ENTRYPOINT and CMD. There is no shell: to expand variables, use "command": "bash" with "args": ["-c", "…"]. |
args |
list of strings | [] |
The command's arguments. Without command, they replace the image's CMD. |
env |
map of string to string | {} |
Environment variables. Yours override what Astraeus injects; see Inside a worker. |
user |
string | the image's USER, else root |
name, uid or uid:gid. |
requested_resources |
object | required | CPU, memory, GPUs, and where the workers may go. |
restart_policy |
OnFailure, Always or Never |
OnFailure |
Whether a worker that ended is started again. lifetime: Service forces Always; on_failure: FailJob forces Never. |
time_limit_seconds |
integer, 60 to 31 622 400 | none | Walltime, counted from when the worker is Running. Past it the worker is stopped and fails; it is never restarted. A limit also lets the run backfill. |
volumes |
list of volumes | [] |
Scratch directories and host paths. |
datavolume_refs |
list of drive mounts | [] |
Drives to mount. |
secret_refs |
list of credential references | [] |
Credentials, as files or environment variables. |
configs |
list of {mounts, value} |
[] |
Literal files: value is written to every path in mounts, read-only. |
security_context |
object | — | Privileges, namespaces, capabilities, ulimits. |
health_check |
object | — | A readiness probe. |
metrics_endpoint |
{port, path, interval_seconds} |
— | A Prometheus endpoint the machine scrapes: port 1–65535, path starting with /, interval_seconds at least 5 (default 15). |
external_accesses |
list | [] |
Ports reachable from outside the cluster. See External access. |
heartbeat_ttl_seconds |
integer, 10 to 86 400 | 60 |
How long a running worker may go without a heartbeat before it turns Stale, then Down. |
network |
{egress} |
— | What the workers may reach outbound. See network.egress. |
isolation |
standard or sandbox |
standard |
sandbox runs each worker under gVisor. Not with GPUs, privileged, or the host's PID or IPC namespace. Placed only on machines that have it. |
identity |
boolean | false |
Give each worker an X.509 certificate and a signed token in /var/run/astraeus/identity. |
outputs |
{answer_path} |
— | Collect up to 64 key=value pairs (4 KiB in all) that the worker writes to $ASTRAEUS_OUTPUTS. |
tools, agent_sandbox |
object | — | Agent runs: the tool policy and the NVIDIA OpenShell sandbox. See Anemoi. |
interactive |
object | — | A server reached through the console, such as a notebook. See Hesperus. |
init_containers |
list | [] |
Validated (image is required) but not run by machines in this version. |
context_bindings |
list | [] |
Validated but not mounted by machines in this version. |
requested_resources#
Give CPU and memory in exactly one of three forms: per worker (cpu_cores and memory_bytes), per GPU (per_gpu), or as a share of a machine (node_fraction). An ask with none of them is refused.
| Field | Type | Default | Description |
|---|---|---|---|
cpu_cores |
integer ≥ 0 | 0 |
Cores per worker. Must be > 0 together with memory_bytes. A hard quota unless cpu_burst is set. |
memory_bytes |
integer ≥ 0 | 0 |
Memory per worker, in bytes. A hard limit: past it the worker is killed (OOMKilled). |
per_gpu |
{cpu_cores, memory_bytes} |
— | Cores and memory for each GPU a worker gets, both > 0. A worker of 8 GPUs gets 8 times as much, so the run's total follows the shape the scheduler picks. Needs GPUs; excludes the other two forms. |
node_fraction |
number in (0, 1] | 0 |
A share of one machine's cores and memory, reserved by placement. The container gets no CPU or memory limit. |
cpu_burst |
boolean | false |
The cores become a weight rather than a ceiling: the worker may use idle cores. Needs cpu_cores or per_gpu.cpu_cores. Memory stays a hard limit. |
gpu_requests |
object | no GPUs | How many GPUs, and which. |
machines |
integer ≥ 0 | 0 |
Run on exactly this many machines, one worker on each, with the GPUs split evenly ("16 GPUs on 2 machines"). 0: the scheduler decides. |
topology |
object | — | Rules on where workers land relative to each other. |
network |
object | auto |
The slowest link allowed between workers. |
node_selection |
object | — | Which machines, by name or label, and how many workers. |
placement |
Any, TightestNetwork, SameRack, SpreadRacks, SpreadFabrics or SameNvlinkDomain |
Any |
The older form of topology. SameRack is keep_within: rack, SameNvlinkDomain is keep_within: nvlink_domain, SpreadRacks and SpreadFabrics are spread: rack and spread: fabric. TightestNetwork and Any add nothing. Where both say something, topology wins. |
Contradictions are refused when the run is created. For example: machines that does not divide gpu_requests.count; machines together with node_selection.count or topology.min_nodes; interconnect: nvlink without GPUs; keep_within: node with a spread; min_nodes above max_nodes; patience_seconds over a week.
gpu_requests#
| Field | Type | Default | Description |
|---|---|---|---|
count |
integer | 0 |
GPUs for the whole run, which the scheduler splits across workers. In an array, per copy. |
models |
list of strings | any | Only GPUs whose model name contains one of these, ignoring case: ["H100", "H200"], ["MI300X"]. |
min_memory_gb |
integer | any | Only GPUs with at least this much memory, in GB. Allows 0.5 GB of slack: an "80 GB" H100 reports 79.6. |
vendor |
string | any | nvidia, amd or intel (apple on a Mac). A GPU that reports no vendor counts as NVIDIA. |
healthy_only |
boolean | false |
Never a GPU at risk (growing memory errors, PCIe replays, NVLink errors, thermal throttling). Without it, such GPUs are used only when no healthy one is free. |
min_per_machine |
integer ≥ 0 | 0 |
When the scheduler picks the shape: never fewer than this many GPUs on one machine. |
policy |
BestEffort, PerNode or All |
BestEffort |
BestEffort: count GPUs, split as the scheduler or machines decides. PerNode: a power of two per worker, one worker per machine. All: every free GPU of each machine it lands on, one worker per GPU machine. |
ids |
list of strings | [] |
Pin exact GPU ids. All of them must be free. |
topology#
These are rules, never broken: when one cannot be met, the run waits and says why.
| Field | Type | Default | Description |
|---|---|---|---|
keep_within |
node, nvlink_domain, fabric or rack |
— | The whole run inside one machine, NVLink domain, InfiniBand fabric or rack. |
spread |
node, rack or fabric |
— | One worker per machine, rack or fabric. |
max_nodes |
integer | 0 (no limit) |
At most this many machines. |
min_nodes |
integer | 0 |
At least this many machines, one worker on each. With max_nodes, a range the scheduler picks within. |
patience_seconds |
integer, at most 604 800 | 0 |
A gang that could start now on a looser network waits up to this long for a tighter one. |
Racks and fabrics are the machine labels topology.astraeus.io/rack and topology.astraeus.io/fabric; NVLink domains are what the GPUs report. See GPUs and placement.
network#
| Field | Type | Default | Description |
|---|---|---|---|
interconnect |
auto, rdma, infiniband or nvlink |
auto |
auto: the best there is room on, Ethernet last. rdma: InfiniBand or RoCE, never plain Ethernet; each worker gets the machine's RDMA devices. infiniband: InfiniBand, not RoCE. nvlink: one machine or one NVLink domain; needs GPUs. |
min_gbps |
integer | 0 |
RDMA ports at least this fast, in Gb/s. Implies RDMA. |
node_selection#
| Field | Type | Default | Description |
|---|---|---|---|
mode |
Exact, Any or All |
Exact |
Exact: workers never share a machine; with names and no count, one worker per named machine. Any: any of names (or any machine); workers may share one. All: one worker on every machine that could take it when the run is created. |
names |
list of strings | [] |
Candidate machines. |
count |
integer, or a numeric string | 0 |
How many workers (what other systems call replicas). With Exact, on distinct machines. Not with machines. |
match_labels |
map of string to string | {} |
Only machines carrying every one of these labels: pool, topology.astraeus.io/site, and so on. The workspace's pools are added to it. |
volumes#
| Field | Type | Default | Description |
|---|---|---|---|
kind |
EmptyDir or Bind |
EmptyDir |
EmptyDir: an empty, writable directory on the machine, removed with the worker's container. Bind: a path of the machine. |
mount_path |
absolute path | required | Where it appears in the container. |
host_path |
absolute path | — | For Bind. In a workspace, only paths the workspace is granted. |
mode |
ReadOnly or ReadWrite |
ReadOnly |
Applies to Bind only. |
datavolume_refs#
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | required | The drive. |
mount_path |
absolute path | the drive's own | Where it appears in the container. |
sub_path |
string | — | Mount this subdirectory of the drive. |
mode |
ReadOnly or ReadWrite |
the drive's | Can only narrow: a read-only drive is never mounted read-write. |
source_index |
integer ≥ 0 | every source | One source of a drive with several. |
secret_refs#
| Field | Type | Default | Description |
|---|---|---|---|
external_secret_name |
string | required | The credential. |
mount_path |
absolute path | — | Mount its keys as files, one file per key, read-only. |
env_vars |
map of key to variable name | {} |
Put key k in the variable env_vars[k]. These override env. |
The machine fetches the value itself. The worker stays in Preparing until the value is there.
security_context#
In a workspace, privileged, pid_mode: host, ipc_mode: host, a non-empty cap_add and a non-empty security_opt need the workspace to be granted privileged work; otherwise the run is refused.
| Field | Type | Default | Description |
|---|---|---|---|
privileged |
boolean | false |
Every capability, every host device, a writable /sys. |
ipc_mode |
host or private |
private | host also gives the machine's /dev/shm. |
pid_mode |
host or private |
private | |
cap_add, cap_drop |
list of strings | [] |
Linux capabilities; the CAP_ prefix is optional. |
security_opt |
list of strings | [] |
Passed to the container runtime as given. |
ulimits |
list of {name, soft, hard} |
see Limits | name without RLIMIT_ (memlock, nofile, stack); -1 is unlimited. |
health_check#
A readiness probe. While it fails, the worker is not routable (endpoints, depends_on with Healthy). A failing probe never restarts the worker.
| Field | Type | Default | Description |
|---|---|---|---|
type |
HTTP, TCP or Exec |
required | HTTP: a GET answered with 200–399. TCP: a connection opens. Exec: a command exits 0. |
path |
string | — | Required for HTTP. |
port |
integer, 1–65535 | — | Required for HTTP and TCP. |
command |
list of strings | — | Required for Exec. |
interval_seconds |
integer | 10 |
Between probes. |
timeout_seconds |
integer | 5 |
Must be less than interval_seconds. |
failure_threshold |
integer | 3 |
Consecutive failures before the worker is unhealthy. |
initial_delay_seconds |
integer | 0 |
Before the first probe. |
network.egress#
| Field | Type | Default | Description |
|---|---|---|---|
default |
allow or deny |
allow |
deny: nothing outbound except allow. |
allow[].host |
string | — | A name the workers may resolve and reach. *.example.com matches one label under it, not example.com itself. |
allow[].cidr |
string | — | An IPv4 network or address. |
allow[].ports |
list of integers | any | TCP and UDP ports. |
cluster |
none, namespace or all |
namespace under deny, all under allow |
Which other workers it may reach. all is refused in a workspace. |
A run with a restricting policy is placed only on machines that can enforce it: the containerd runtime, with workers on the node network. It cannot be combined with privileged, cap_add: [NET_ADMIN] (or ALL) or pid_mode: host.
What the cluster returns#
POST /jobs answers 201 Created with the run as stored: metadata and spec, defaults applied. GET /jobs/<name> adds:
| Field | Description |
|---|---|
status |
state, reason, created_at, updated_at, started_at, finished_at. See Run and worker states. |
history |
Every state change, oldest first: state, reason, transition_at. |
task_names, task_counts |
The workers, and how many are in each state. |
leader_task |
The leader: rank 0 of the leader group. |
expected_task_count |
How many workers the run has. |
shape |
When the scheduler picked the machines: {tasks, gpus_each, why, chosen_at}, for example why: "2 machines × 8 GPUs on InfiniBand fabric ib-0". |
Errors have the shape {"code", "message", "detail"}. A validation failure is 400 VALIDATION_ERROR with one detail entry per field:
{"code": "VALIDATION_ERROR", "message": "Failed to validate object",
"detail": {"task_template.time_limit_seconds": "Must be between 60 and 31622400 seconds, got 30"}}
See Errors for every code.