Skip to content

Run specification#

A run is the object you send to POST …/jobs under a workspace's cluster path (see REST API); the console's New run form and astra build it for you. This page lists every field Astraeus accepts, with its type, default and limits, and what it does. Use it when you write a run by hand, or when Astraeus refuses one and names a field.

In the API a run is a job and a worker is a task. The field names keep those words.

Shape of a run#

run.json
{
  "metadata": { "name": "train-resnet", "labels": { "team": "vision" } },
  "spec": {
    "task_template": {
      "image": "nvcr.io/nvidia/pytorch:24.08-py3",
      "command": "python",
      "args": ["train.py", "--epochs", "90"],
      "requested_resources": {
        "cpu_cores": 8,
        "memory_bytes": 68719476736,
        "gpu_requests": { "count": 1 }
      }
    }
  }
}
  • metadata names the run. spec says what it runs (task_template, or several task_groups) and how its workers start, fail and end.
  • The API reads JSON. If you keep specs in YAML, convert them as you send them (yq -o=json run.yaml). See Submit a run.
  • Unknown fields are refused, not ignored. A typo such as spec.replicas or gpu_requests.model returns 400 UNKNOWN_FIELD and names the field. The check covers spec and spec.task_template.requested_resources.
  • Sizes are integers in bytes; durations are integers in seconds.

A complete example#

The run below uses most options. Each numbered marker explains its line. Convert it to JSON before you send it.

run.yaml
metadata:
  name: llama-finetune            # (1)!
  labels: { team: nlp }
spec:
  start: Gang                     # (2)!
  on_failure: RestartJob
  lifetime: Batch
  priority: 10                    # (3)!
  task_template:
    image: registry.example.com/nlp/train:2026-09
    pull_policy: IfNotPresent
    registry_secret_ref: { external_secret_name: registry-login }   # (4)!
    command: torchrun
    args: ["--nproc-per-node=8", "train.py", "--config", "configs/70b.yaml"]
    env: { NCCL_DEBUG: INFO, HF_HOME: /data/hf }
    restart_policy: OnFailure
    time_limit_seconds: 86400     # (5)!
    requested_resources:
      gpu_requests:
        count: 16                 # (6)!
        models: [H100, H200]
        min_memory_gb: 80
        healthy_only: true
      machines: 2                 # (7)!
      per_gpu: { cpu_cores: 12, memory_bytes: 128849018880 }
      network: { interconnect: infiniband, min_gbps: 200 }
      topology: { keep_within: rack, patience_seconds: 600 }
      node_selection:
        mode: Any
        match_labels: { pool: h100 }
    datavolume_refs:
      - { name: datasets, mount_path: /data, mode: ReadOnly }
      - { name: checkpoints, mount_path: /ckpt, mode: ReadWrite }
    secret_refs:
      - external_secret_name: hf-token
        env_vars: { token: HF_TOKEN }
    security_context:
      ulimits: [{ name: stack, soft: 67108864, hard: 67108864 }]
  1. Letters, digits, - and _; at most 57 characters. The workers are named llama-finetune-0 and llama-finetune-1.
  2. Every worker is placed and started together. A failure restarts all of them.
  3. At most the workspace's maximum priority, which is 0 unless an admin raised it.
  4. A credential whose value is JSON: {"username": …, "password": …, "server": …}.
  5. The workers are stopped and the run fails after 24 hours of running. A time limit also lets the run backfill.
  6. The run's total. With machines: 2, each of the two machines gets 8.
  7. Leave it out to let the scheduler pick the fewest machines on the best network.

Minimal examples#

{"metadata": {"name": "export"},
 "spec": {"task_template": {
   "image": "python:3.12", "command": "python", "args": ["export.py"],
   "requested_resources": {"cpu_cores": 4, "memory_bytes": 17179869184,
                           "gpu_requests": {"count": 1}}}}}
{"metadata": {"name": "etl"},
 "spec": {"task_template": {
   "image": "busybox", "command": "sh", "args": ["-c", "echo hello"],
   "requested_resources": {"cpu_cores": 1, "memory_bytes": 268435456}}}}
{"metadata": {"name": "pretrain"},
 "spec": {"start": "Gang", "task_template": {
   "image": "nvcr.io/nvidia/pytorch:24.08-py3", "command": "bash",
   "args": ["-c", "torchrun --nnodes=$WORLD_SIZE --node-rank=$RANK --nproc-per-node=$ASTRAEUS_GPU_COUNT --master-addr=$MASTER_ADDR --master-port=$MASTER_PORT train.py"],
   "requested_resources": {"gpu_requests": {"count": 16},
                           "per_gpu": {"cpu_cores": 12, "memory_bytes": 128849018880}}}}}
{"metadata": {"name": "sweep"},
 "spec": {"array": {"size": 100, "max_parallel": 10}, "task_template": {
   "image": "python:3.12", "command": "python", "args": ["sweep.py"],
   "requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296}}}}

metadata#

Field Type Default Description
name string required The run's name, unique in its workspace on that cluster. Letters, digits, - and _; no - at either end; no .. At most 57 characters; with task_groups, at most 57 minus the length of the longest group name minus 1, so that every <run>-<group>-<rank> stays a DNS label. Lowercase is recommended: the console only accepts [a-z0-9-].
labels map of string to string {} Your labels. Copied onto every worker. Keys under astraeus.io/ are reserved and refused.

The cluster adds astraeus.io/created-by (who submitted the run). Workers also carry job.astraeus.io/name, job.astraeus.io/replica-index, job.astraeus.io/group and job.astraeus.io/group-replica-index.

spec#

Field Type Default Description
task_template worker template — What every worker runs. Required unless task_groups is set; must be left out when it is.
task_groups list of worker groups [] Several kinds of worker in one run: a launcher and its workers, an engine and its clients.
leader_group string the first group With task_groups: the group whose rank-0 worker leads (its machine is MASTER_ADDR).
complete_with string — With task_groups: the group whose completion completes the run. The other groups are then stopped, and one of them failing for good fails the run. Not with lifetime: Service.
start Gang, MinAvailable or Independent see defaults How workers are admitted and started. Gang: every worker placed and started together, or none. MinAvailable: starts once min_available workers can be placed together; the rest join as room appears. Independent: each worker when it fits.
min_available integer ≥ 1 — Required with start: MinAvailable; refused otherwise.
on_failure RestartJob, RestartTask or FailJob see defaults What a worker's failure does. RestartJob: restart every worker. RestartTask: restart the one that failed; in a gang, the leader's failure still restarts all. FailJob: no restart; the run fails and its other workers are stopped.
lifetime Batch or Service Batch Batch runs to completion. Service runs until you delete it: every worker is restarted whenever it ends (restart_policy becomes Always), and the run never completes.
policy Gang, GangIndependent or Independent Independent The older form of start and on_failure, still accepted. When you set either of those, the cluster rewrites policy to match.
priority integer, −1000 to 1000 0 Place in the queue: higher first, and may preempt lower. Capped by the workspace's maximum priority. See Priorities.
array object — Run the template size times as independent copies.
model object — A model of the workspace mounted in the workers.
shared_volumes list [] Stored and validated, but not provisioned in this version: no drive is created. Use a drive and datavolume_refs.
output_volumes list [] As shared_volumes, with retain defaulting to true. Not provisioned in this version.
idle_auto_delete object — Stored and validated (enabled, idle_timeout, grace_period_at_start, warn_before_kill, thresholds, bypass_during_hours as "HH:MM-HH:MM <IANA zone>"), but nothing acts on it in this version.
retry_on_heartbeat_expired boolean false Accepted for compatibility; not read. A worker lost with its machine is restarted according to its restart_policy.
max_heartbeat_retries integer, 0 to 1000 0 Accepted for compatibility; not read.

Defaults for start and on_failure#

You set start becomes on_failure becomes
neither, and no policy Independent RestartTask
start: Gang only Gang RestartJob
start: MinAvailable or Independent only as set RestartTask
policy: Gang Gang RestartJob
policy: GangIndependent Gang RestartTask

The API and the tools differ

The API's default is Independent. The console sets Gang for you when the run may take several machines (more than one GPU and no array, or more than one machine). astra astraeus run and astra slurm sbatch do not set start, so their runs start Independent. In specs you send to the API, set start yourself.

Rules checked when a run is created#

  • lifetime: Service cannot have an array or on_failure: FailJob.
  • start: Gang cannot have task_groups with depends_on: a gang starts every worker together.
  • An array needs start: Independent (or no start) and cannot have task_groups.
  • Group names are unique. depends_on, leader_group and complete_with name groups that exist. A group cannot depend on itself.
  • An enabled external access name is unique across the run's groups.

array#

Field Type Default Description
size integer, 1 to 10 000 required How many workers. Each has the template's whole ask. Index i gets ASTRAEUS_ARRAY_TASK_ID=i (from 0) and ASTRAEUS_ARRAY_SIZE.
max_parallel integer, 0 to size 0 At most this many hold machines at once. 0: no cap.

See Arrays.

model#

Field Type Default Description
name string required A model of the workspace.
mount_path absolute path /model Where its weights appear, read-only. ASTRAEUS_MODEL_PATH names it, and ASTRAEUS_MODEL_FILE names a single GGUF file.
engine boolean false Also serve the model beside the run: an engine worker group serves it, your workers become the main group with OPENAI_BASE_URL and OPENAI_MODEL set, and the run completes when main does. Not with lifetime: Service or an array.

A cluster that does not serve models refuses a run with model (400 MODELS_NOT_SERVED). See Eos.

task_groups: worker groups#

Field Type Default Description
name string required The group's name. Its workers are named <run>-<name>-<i>.
task_template worker template required What this group's workers run.
depends_on list of {group_name, condition} [] Groups that must be up before this one is placed. condition: Running (the default) or Healthy (running and passing its health_check). Not with start: Gang.

Ranks are global across groups, in the order the groups are listed. With a group worker of two workers and a group ps of one, the workers are run-worker-0 (rank 0), run-worker-1 (rank 1) and run-ps-0 (rank 2).

task_template: the worker template#

Field Type Default Description
image string required The container image. busybox means docker.io/library/busybox:latest.
pull_policy IfNotPresent, Always or None IfNotPresent When the machine pulls the image. None: never; the worker fails if the image is not already on the machine.
registry_secret_ref {external_secret_name} — A credential whose value is JSON {"username", "password", "server"}, used to pull from a private registry.
command string the image's One executable. It replaces the image's ENTRYPOINT and CMD. There is no shell: to expand variables, use "command": "bash" with "args": ["-c", "…"].
args list of strings [] The command's arguments. Without command, they replace the image's CMD.
env map of string to string {} Environment variables. Yours override what Astraeus injects; see Inside a worker.
user string the image's USER, else root name, uid or uid:gid.
requested_resources object required CPU, memory, GPUs, and where the workers may go.
restart_policy OnFailure, Always or Never OnFailure Whether a worker that ended is started again. lifetime: Service forces Always; on_failure: FailJob forces Never.
time_limit_seconds integer, 60 to 31 622 400 none Walltime, counted from when the worker is Running. Past it the worker is stopped and fails; it is never restarted. A limit also lets the run backfill.
volumes list of volumes [] Scratch directories and host paths.
datavolume_refs list of drive mounts [] Drives to mount.
secret_refs list of credential references [] Credentials, as files or environment variables.
configs list of {mounts, value} [] Literal files: value is written to every path in mounts, read-only.
security_context object — Privileges, namespaces, capabilities, ulimits.
health_check object — A readiness probe.
metrics_endpoint {port, path, interval_seconds} — A Prometheus endpoint the machine scrapes: port 1–65535, path starting with /, interval_seconds at least 5 (default 15).
external_accesses list [] Ports reachable from outside the cluster. See External access.
heartbeat_ttl_seconds integer, 10 to 86 400 60 How long a running worker may go without a heartbeat before it turns Stale, then Down.
network {egress} — What the workers may reach outbound. See network.egress.
isolation standard or sandbox standard sandbox runs each worker under gVisor. Not with GPUs, privileged, or the host's PID or IPC namespace. Placed only on machines that have it.
identity boolean false Give each worker an X.509 certificate and a signed token in /var/run/astraeus/identity.
outputs {answer_path} — Collect up to 64 key=value pairs (4 KiB in all) that the worker writes to $ASTRAEUS_OUTPUTS.
tools, agent_sandbox object — Agent runs: the tool policy and the NVIDIA OpenShell sandbox. See Anemoi.
interactive object — A server reached through the console, such as a notebook. See Hesperus.
init_containers list [] Validated (image is required) but not run by machines in this version.
context_bindings list [] Validated but not mounted by machines in this version.

requested_resources#

Give CPU and memory in exactly one of three forms: per worker (cpu_cores and memory_bytes), per GPU (per_gpu), or as a share of a machine (node_fraction). An ask with none of them is refused.

Field Type Default Description
cpu_cores integer ≥ 0 0 Cores per worker. Must be > 0 together with memory_bytes. A hard quota unless cpu_burst is set.
memory_bytes integer ≥ 0 0 Memory per worker, in bytes. A hard limit: past it the worker is killed (OOMKilled).
per_gpu {cpu_cores, memory_bytes} — Cores and memory for each GPU a worker gets, both > 0. A worker of 8 GPUs gets 8 times as much, so the run's total follows the shape the scheduler picks. Needs GPUs; excludes the other two forms.
node_fraction number in (0, 1] 0 A share of one machine's cores and memory, reserved by placement. The container gets no CPU or memory limit.
cpu_burst boolean false The cores become a weight rather than a ceiling: the worker may use idle cores. Needs cpu_cores or per_gpu.cpu_cores. Memory stays a hard limit.
gpu_requests object no GPUs How many GPUs, and which.
machines integer ≥ 0 0 Run on exactly this many machines, one worker on each, with the GPUs split evenly ("16 GPUs on 2 machines"). 0: the scheduler decides.
topology object — Rules on where workers land relative to each other.
network object auto The slowest link allowed between workers.
node_selection object — Which machines, by name or label, and how many workers.
placement Any, TightestNetwork, SameRack, SpreadRacks, SpreadFabrics or SameNvlinkDomain Any The older form of topology. SameRack is keep_within: rack, SameNvlinkDomain is keep_within: nvlink_domain, SpreadRacks and SpreadFabrics are spread: rack and spread: fabric. TightestNetwork and Any add nothing. Where both say something, topology wins.

Contradictions are refused when the run is created. For example: machines that does not divide gpu_requests.count; machines together with node_selection.count or topology.min_nodes; interconnect: nvlink without GPUs; keep_within: node with a spread; min_nodes above max_nodes; patience_seconds over a week.

gpu_requests#

Field Type Default Description
count integer 0 GPUs for the whole run, which the scheduler splits across workers. In an array, per copy.
models list of strings any Only GPUs whose model name contains one of these, ignoring case: ["H100", "H200"], ["MI300X"].
min_memory_gb integer any Only GPUs with at least this much memory, in GB. Allows 0.5 GB of slack: an "80 GB" H100 reports 79.6.
vendor string any nvidia, amd or intel (apple on a Mac). A GPU that reports no vendor counts as NVIDIA.
healthy_only boolean false Never a GPU at risk (growing memory errors, PCIe replays, NVLink errors, thermal throttling). Without it, such GPUs are used only when no healthy one is free.
min_per_machine integer ≥ 0 0 When the scheduler picks the shape: never fewer than this many GPUs on one machine.
policy BestEffort, PerNode or All BestEffort BestEffort: count GPUs, split as the scheduler or machines decides. PerNode: a power of two per worker, one worker per machine. All: every free GPU of each machine it lands on, one worker per GPU machine.
ids list of strings [] Pin exact GPU ids. All of them must be free.

topology#

These are rules, never broken: when one cannot be met, the run waits and says why.

Field Type Default Description
keep_within node, nvlink_domain, fabric or rack — The whole run inside one machine, NVLink domain, InfiniBand fabric or rack.
spread node, rack or fabric — One worker per machine, rack or fabric.
max_nodes integer 0 (no limit) At most this many machines.
min_nodes integer 0 At least this many machines, one worker on each. With max_nodes, a range the scheduler picks within.
patience_seconds integer, at most 604 800 0 A gang that could start now on a looser network waits up to this long for a tighter one.

Racks and fabrics are the machine labels topology.astraeus.io/rack and topology.astraeus.io/fabric; NVLink domains are what the GPUs report. See GPUs and placement.

network#

Field Type Default Description
interconnect auto, rdma, infiniband or nvlink auto auto: the best there is room on, Ethernet last. rdma: InfiniBand or RoCE, never plain Ethernet; each worker gets the machine's RDMA devices. infiniband: InfiniBand, not RoCE. nvlink: one machine or one NVLink domain; needs GPUs.
min_gbps integer 0 RDMA ports at least this fast, in Gb/s. Implies RDMA.

node_selection#

Field Type Default Description
mode Exact, Any or All Exact Exact: workers never share a machine; with names and no count, one worker per named machine. Any: any of names (or any machine); workers may share one. All: one worker on every machine that could take it when the run is created.
names list of strings [] Candidate machines.
count integer, or a numeric string 0 How many workers (what other systems call replicas). With Exact, on distinct machines. Not with machines.
match_labels map of string to string {} Only machines carrying every one of these labels: pool, topology.astraeus.io/site, and so on. The workspace's pools are added to it.

volumes#

Field Type Default Description
kind EmptyDir or Bind EmptyDir EmptyDir: an empty, writable directory on the machine, removed with the worker's container. Bind: a path of the machine.
mount_path absolute path required Where it appears in the container.
host_path absolute path — For Bind. In a workspace, only paths the workspace is granted.
mode ReadOnly or ReadWrite ReadOnly Applies to Bind only.

datavolume_refs#

Field Type Default Description
name string required The drive.
mount_path absolute path the drive's own Where it appears in the container.
sub_path string — Mount this subdirectory of the drive.
mode ReadOnly or ReadWrite the drive's Can only narrow: a read-only drive is never mounted read-write.
source_index integer ≥ 0 every source One source of a drive with several.

secret_refs#

Field Type Default Description
external_secret_name string required The credential.
mount_path absolute path — Mount its keys as files, one file per key, read-only.
env_vars map of key to variable name {} Put key k in the variable env_vars[k]. These override env.

The machine fetches the value itself. The worker stays in Preparing until the value is there.

security_context#

In a workspace, privileged, pid_mode: host, ipc_mode: host, a non-empty cap_add and a non-empty security_opt need the workspace to be granted privileged work; otherwise the run is refused.

Field Type Default Description
privileged boolean false Every capability, every host device, a writable /sys.
ipc_mode host or private private host also gives the machine's /dev/shm.
pid_mode host or private private
cap_add, cap_drop list of strings [] Linux capabilities; the CAP_ prefix is optional.
security_opt list of strings [] Passed to the container runtime as given.
ulimits list of {name, soft, hard} see Limits name without RLIMIT_ (memlock, nofile, stack); -1 is unlimited.

health_check#

A readiness probe. While it fails, the worker is not routable (endpoints, depends_on with Healthy). A failing probe never restarts the worker.

Field Type Default Description
type HTTP, TCP or Exec required HTTP: a GET answered with 200–399. TCP: a connection opens. Exec: a command exits 0.
path string — Required for HTTP.
port integer, 1–65535 — Required for HTTP and TCP.
command list of strings — Required for Exec.
interval_seconds integer 10 Between probes.
timeout_seconds integer 5 Must be less than interval_seconds.
failure_threshold integer 3 Consecutive failures before the worker is unhealthy.
initial_delay_seconds integer 0 Before the first probe.

network.egress#

Field Type Default Description
default allow or deny allow deny: nothing outbound except allow.
allow[].host string — A name the workers may resolve and reach. *.example.com matches one label under it, not example.com itself.
allow[].cidr string — An IPv4 network or address.
allow[].ports list of integers any TCP and UDP ports.
cluster none, namespace or all namespace under deny, all under allow Which other workers it may reach. all is refused in a workspace.

A run with a restricting policy is placed only on machines that can enforce it: the containerd runtime, with workers on the node network. It cannot be combined with privileged, cap_add: [NET_ADMIN] (or ALL) or pid_mode: host.

What the cluster returns#

POST /jobs answers 201 Created with the run as stored: metadata and spec, defaults applied. GET /jobs/<name> adds:

Field Description
status state, reason, created_at, updated_at, started_at, finished_at. See Run and worker states.
history Every state change, oldest first: state, reason, transition_at.
task_names, task_counts The workers, and how many are in each state.
leader_task The leader: rank 0 of the leader group.
expected_task_count How many workers the run has.
shape When the scheduler picked the machines: {tasks, gpus_each, why, chosen_at}, for example why: "2 machines × 8 GPUs on InfiniBand fabric ib-0".

Errors have the shape {"code", "message", "detail"}. A validation failure is 400 VALIDATION_ERROR with one detail entry per field:

{"code": "VALIDATION_ERROR", "message": "Failed to validate object",
 "detail": {"task_template.time_limit_seconds": "Must be between 60 and 31622400 seconds, got 30"}}

See Errors for every code.