Skip to content

Run and worker states#

A run's state is derived from its workers' states, and every change is recorded with a reason. This page lists every state, how each one is reached, and the exact reasons Astraeus writes, so you can read a run's status, its history, or an event stream without guessing.

The API calls a run a job and a worker a task. Both have a status (state, reason, timestamps) and a history of every change, oldest first. Events are derived from that history.

Run states#

State Meaning
Pending No worker is running yet: the run waits for room, for its workers to be created, or to be placed again after a preemption. The reason says what it waits for.
Starting Gang runs only: the whole gang is placed and its workers are preparing, waiting at the barrier or starting. Reason: Starting together: 3 of 4 ready.
Running At least one worker runs (for a gang, every worker). Reason: 2 of 2 workers running (serving for a service).
Completed Every worker exited 0, or the complete_with group completed. Terminal.
Failed A worker failed for good under the run's policy. Terminal. The reason names the worker, the machine and why.
Cancelled Every worker ended and at least one was cancelled (stopped by a person or the API) rather than completed. Terminal.

Terminal runs keep nothing running: any worker still active is stopped with the reason Run failed, Run completed or Run cancelled.

stateDiagram-v2
    [*] --> Pending: Run created
    Pending --> Starting: gang placed
    Starting --> Running: every member runs
    Pending --> Running: a worker runs
    Running --> Pending: preempted, or every worker restarting
    Starting --> Pending: gang restart
    Running --> Completed
    Running --> Failed
    Starting --> Failed
    Pending --> Failed
    Running --> Cancelled
    Pending --> Cancelled
    Completed --> [*]
    Failed --> [*]
    Cancelled --> [*]

A deleted run is not a state: deleting removes the run, its workers and their logs. The console's runs list can still show it as Deleted with Include deleted.

Run reasons#

Reason State Meaning
Run created Pending Accepted; the scheduler has not looked yet.
Waiting for workers to be created Pending The scheduler is materialising the workers.
Cannot plan workers: <why> Pending The run cannot be split into workers on this cluster yet: cannot evenly distribute 12 gpus across 5 machines, 16 gpus across 2 machines is 8 per worker, not a power of 2, no way to run 16 GPUs on these machines: 4 free on the largest one now, no machine can host a worker of group <g> right now. Looked at again every 60 s.
Waiting for workers to start Pending Workers exist and are not placed; usually replaced by the first waiting worker's reason (see pending reasons).
<reason of the first waiting worker> Pending The run shows why its first waiting worker waits.
<cause>; waiting to run again Pending Preempted: Preempted by higher-priority work (urgent-eval); waiting to run again.
Starting together: <r> of <n> ready Starting Gang members at the barrier or beyond.
<r> of <n> workers running Running serving instead of running for lifetime: Service.
All workers completed Completed
Worker group <g> completed Completed Runs with complete_with.
Worker <w> failed on <machine>: <reason> Failed One worker failed for good.
Worker <w> was lost on <machine>: <reason> Failed Its machine stopped answering and it will not be restarted.
… ; the gang does not run without it Failed A gang member failed for good.
<k> of <n> workers failed; the first: … Failed Several failed.
Worker <w>: Exceeded its time limit of <d> Failed A worker ran past time_limit_seconds.
… ; worker group <g> cannot finish without it Failed A serving group failed for good in a complete_with run.
All workers ended; some were cancelled Cancelled
Quota reached: <why> any A history entry (and an event) written the first time the run waits on its workspace's quota.

Worker states#

State Meaning
Pending Not running and waiting to be placed (or placed, and its machine has not picked it up yet).
Preparing Its machine is setting it up: credentials, drives, network, identity.
Pulling Its machine is pulling the image.
ReadyToStart Gang runs: set up and waiting at the barrier for every member.
Starting The container was started.
Running Its machine reports it running and sends heartbeats.
Stale Running, but its heartbeat is late.
Stopping Asked to stop (cancelled, preempted, out of time); its machine is stopping the container.
Completed Exited 0. Terminal unless restart_policy: Always.
Failed Exited non-zero, could not start, or ran out of restarts. Terminal when it will not be restarted.
Cancelled Stopped before or while running. Terminal, except after a preemption (it is queued again).
Down Lost: its machine stopped answering, or its heartbeat expired. Restarted under its restart_policy.
stateDiagram-v2
    [*] --> Pending
    Pending --> Preparing: placed, machine picks it up
    Preparing --> Pulling
    Preparing --> Starting
    Pulling --> Starting
    Preparing --> ReadyToStart: gang
    Pulling --> ReadyToStart: gang
    ReadyToStart --> Starting: every member ready
    Starting --> Running
    Running --> Stale: heartbeat late
    Stale --> Running: heartbeat back
    Stale --> Down: no heartbeat for its TTL
    Running --> Down: machine lost
    Running --> Completed: exit 0
    Running --> Failed: exit ≠ 0, OOM
    Preparing --> Failed: image, config, start errors
    Pulling --> Failed
    Running --> Stopping: stop, preemption, time limit
    Stopping --> Cancelled
    Pending --> Cancelled: stopped before placement
    Failed --> Pending: restarted by policy
    Down --> Pending: restarted by policy
    Completed --> Pending: restart_policy Always
    Cancelled --> Pending: preempted, queued again

Liveness#

Each running worker must send a heartbeat within heartbeat_ttl_seconds (default 60 s).

From Condition To Reason
Running no heartbeat (after 15 s of grace) Stale Heartbeat lease expired
Stale heartbeat again, machine up Running Heartbeat received
Stale no heartbeat for its TTL Down No heartbeat for the worker's full TTL after going stale
any active state its machine is Down Down Machine lost: its lease expired
any active state its machine was removed Down Machine removed from the cluster
Stopping its machine is Down or removed Cancelled Machine lost while stopping / Machine removed from the cluster while stopping

Restarts#

A worker that ended is started again when its policy says so:

restart_policy Restarted after
OnFailure (default) Failed or Down
Always Failed, Down or Completed
Never never (Down with Never fails for good)
  • Back-off: 10 s before the first restart, doubling each time, at most 5 min. While it waits, the worker shows when it restarts (restart_at); the console shows Restarting in a loop.
  • Budget: 10 restarts. The next failure is final: Retry budget exhausted after 10/10 restarts.
  • Reset: a worker that stays Running for 10 min has its count reset to 0.
  • Who restarts: under on_failure: RestartJob (the default for gangs) every worker is requeued: Gang restart: train-1 failed. Under RestartTask, only the failed one: Restarted by policy (restart 2/10); in a gang, a leader failure still restarts all.
  • Not counted: a preempted worker is queued again without spending its budget: Preempted by <run>; queued again.

Worker reasons from its machine#

Reason State Meaning
Waiting for the node network (subnet assignment and mesh) Preparing The machine's network is not ready yet.
Waiting for its external port to be assigned Preparing An external access needs its port.
Waiting for data volume <d>: <why> Preparing A drive is not mountable yet.
Waiting for drive <d> to be filled on <machine> Preparing A placed drive's copy is being made.
waiting for external secret <c> (key <k>) to be synced to this node Preparing A credential is not on the machine yet.
waiting for registry secret <c> to be synced to this node Preparing The registry credential is not there yet.
Waiting for the worker's identity: <why> Preparing identity: true and no certificate yet.
Pulling <image> Pulling
Waiting at the gang barrier ReadyToStart
Container started Starting
Container running Running
Exited with code <n> Completed (0) or Failed
OOMKilled: the container exceeded its memory limit Failed Raise memory_bytes (or per_gpu.memory_bytes).
ErrImagePull: <error> Failed Wrong image name, tag, or registry credential.
Image <image> is not present and the pull policy is None Failed
Container create failed: <error> / Container start failed: <error> Failed The runtime refused the container.
Container not found on the node Failed The container disappeared from the machine.
Invalid task: <error> Failed A mount path or host path the machine refuses.
Writing config <path>: <error> Failed
Outbound policy: <why>, Agent sandbox: <why>, Model server: <why>, Interactive token: <why> Failed The feature could not be set up on the machine.
Stopped on request Cancelled The machine confirmed a stop.

Stop reasons and causes#

A stop moves a placed worker to Stopping (a Pending one straight to Cancelled); its machine sends the stop signal, waits for the grace period, kills it, and confirms Cancelled. Two stops carry a cause that decides what happens next:

Cause Reason Then
Preemption Preempted by higher-priority work (<run>) Queued again once gone (a gang once every member is gone); no restart counted.
Time limit Exceeded its time limit of <d> Fails for good; never restarted.

Other stop reasons: A gang member exhausted its restart budget, A worker failed, and the run fails on a failure, A gang member exceeded its time limit, A worker exceeded its time limit, Gang barrier timed out waiting for all members (requeued after 15 min at the barrier), Run failed / Run completed / Run cancelled, and a reason given to POST /tasks/<w>/stop (default Stopped by scheduler).

Pending reasons#

While a worker waits to be placed, its reason (and its run's) says why. The console turns these into a headline and advice; the table shows the exact text.

Reason Meaning What to do
No machine fits: <r> of <n> machines ready No machine passes every filter; placement.nodes lists each machine's reason. Read the per-machine reasons (how).
Gang cannot be placed whole: <r> of <n> machines ready Not every gang member fits at once. Wait, ask for fewer GPUs, or loosen topology rules.
Cannot place the minimum to start together: … MinAvailable: the first min_available do not fit together. Lower min_available.
… : no machines are registered The workspace can use no machine on this cluster. Add a machine, or check the workspace's pools.
Choose where to keep data on <machine> A placed drive needs a data location on that machine. Choose one on the machine's page.
Waits for namespace quota: gpus 6 in use + 4 asked > 8 Over the workspace's quota (Gang waits for … for a gang). Waiting on quota never blocks others. Wait for the workspace's other work to end, or ask an admin.
Waits for more than the namespace's quota allows at all (…): raise the quota or ask for less The run alone exceeds the quota. Ask for less, or raise the quota.
Waits for worker group <g> to be running / healthy depends_on.
Array runs at most <k> workers at once array.max_parallel.
Waiting up to 6 min more for InfiniBand fabric ib-0 (could start now on Ethernet) topology.patience_seconds.
Next to run (priority <p>): starts by <time> UTC as running work ends The head of the queue; room is held for it.
Next to run (priority <p>): waits for running work to end The head; running work has no time limit, so no start time is known.
Queued behind higher-priority work (<run>) A higher-priority run is the head.
Queued behind <run>, which starts by <time> UTC / Queued behind <run> The head has the same or lower priority and earlier claim; this run cannot fit around it. Give the run a time limit so it can backfill.
Preempting <k> lower-priority worker(s) of <runs> to start This run is the head and is stopping others.
Starting: waiting for <k> preempted worker(s) to stop

Runs from another workspace are never named: such reasons say (a job in another namespace) instead.

Per-machine reasons#

placement.nodes[].reasons (the console's What keeps it out column) uses these:

Reason Meaning
machine is not ready (<state>) / (no capacity reported) Down, or not yet reporting.
under pressure: <condition> Memory or disk pressure: no new work.
machine does not accept runs: <why> / machine is cordoned: <why> An admin stopped new work there.
worker carries none of the labels this machine accepts The machine takes only work selecting its labels.
machine is not among the selected names / machine lacks label <k>=<v> node_selection (pools included).
no active RDMA port / no active InfiniBand port at 200 Gb/s or faster (…) network asked for RDMA.
another worker of this run is already on the machine Workers spread one per machine.
insufficient cpu: need 32 cores, 16 free
insufficient memory: need <bytes> bytes, <bytes> free
insufficient gpus: need 8, 4 free and usable
needs 8 healthy GPUs, 6 free (2 more free but at risk) healthy_only.
no H100 GPU, no nvidia H100 or H200 with at least 80 GB GPU No GPU on it matches gpu_requests.
gpu fault on the machine: GPU 3: Xid 79… A GPU reported a fatal fault.
the run is kept on one machine, <m> / the run is kept within rack <r> / no rack set on this machine (topology.astraeus.io/rack) / the run may use at most <n> machines / rack <r> already has a worker of this run topology rules.
drive <d> does not exist / drive <d>: <why> / drive <d> is on <m>, in another site (<s>) / drive <d> needs <size> at this machine's data location, <size> free Drives.
machine cannot enforce an outbound policy (…), machine cannot run sandboxed workers (…), machine serves no tool gateway to its workers (…) The run needs a capability the machine lacks.

Exit codes#

Exit State Usual cause
0 Completed Success.
1–125 Failed The program's own error.
126, 127 Failed Command not executable, or not found in the image.
137 (128 + 9) Failed Killed with SIGKILL. Out of memory is reported as OOMKilled instead.
143 (128 + 15) Failed Exited on SIGTERM without handling it, outside a stop.

A worker that Astraeus stops (cancel, preemption, time limit) ends Cancelled whatever its exit code; a time-limit stop then fails the run. The last failure's exit code, reason and last 50 lines of log (up to 16 KiB) are kept on the worker as last_failure.