Run and worker states#
A run's state is derived from its workers' states, and every change is recorded with a reason. This page lists every state, how each one is reached, and the exact reasons Astraeus writes, so you can read a run's status, its history, or an event stream without guessing.
The API calls a run a job and a worker a task. Both have a status (state, reason, timestamps) and a history of every change, oldest first. Events are derived from that history.
Run states#
| State | Meaning |
|---|---|
Pending |
No worker is running yet: the run waits for room, for its workers to be created, or to be placed again after a preemption. The reason says what it waits for. |
Starting |
Gang runs only: the whole gang is placed and its workers are preparing, waiting at the barrier or starting. Reason: Starting together: 3 of 4 ready. |
Running |
At least one worker runs (for a gang, every worker). Reason: 2 of 2 workers running (serving for a service). |
Completed |
Every worker exited 0, or the complete_with group completed. Terminal. |
Failed |
A worker failed for good under the run's policy. Terminal. The reason names the worker, the machine and why. |
Cancelled |
Every worker ended and at least one was cancelled (stopped by a person or the API) rather than completed. Terminal. |
Terminal runs keep nothing running: any worker still active is stopped with the reason Run failed, Run completed or Run cancelled.
stateDiagram-v2
[*] --> Pending: Run created
Pending --> Starting: gang placed
Starting --> Running: every member runs
Pending --> Running: a worker runs
Running --> Pending: preempted, or every worker restarting
Starting --> Pending: gang restart
Running --> Completed
Running --> Failed
Starting --> Failed
Pending --> Failed
Running --> Cancelled
Pending --> Cancelled
Completed --> [*]
Failed --> [*]
Cancelled --> [*]
A deleted run is not a state: deleting removes the run, its workers and their logs. The console's runs list can still show it as Deleted with Include deleted.
Run reasons#
| Reason | State | Meaning |
|---|---|---|
Run created |
Pending |
Accepted; the scheduler has not looked yet. |
Waiting for workers to be created |
Pending |
The scheduler is materialising the workers. |
Cannot plan workers: <why> |
Pending |
The run cannot be split into workers on this cluster yet: cannot evenly distribute 12 gpus across 5 machines, 16 gpus across 2 machines is 8 per worker, not a power of 2, no way to run 16 GPUs on these machines: 4 free on the largest one now, no machine can host a worker of group <g> right now. Looked at again every 60 s. |
Waiting for workers to start |
Pending |
Workers exist and are not placed; usually replaced by the first waiting worker's reason (see pending reasons). |
<reason of the first waiting worker> |
Pending |
The run shows why its first waiting worker waits. |
<cause>; waiting to run again |
Pending |
Preempted: Preempted by higher-priority work (urgent-eval); waiting to run again. |
Starting together: <r> of <n> ready |
Starting |
Gang members at the barrier or beyond. |
<r> of <n> workers running |
Running |
serving instead of running for lifetime: Service. |
All workers completed |
Completed |
|
Worker group <g> completed |
Completed |
Runs with complete_with. |
Worker <w> failed on <machine>: <reason> |
Failed |
One worker failed for good. |
Worker <w> was lost on <machine>: <reason> |
Failed |
Its machine stopped answering and it will not be restarted. |
… ; the gang does not run without it |
Failed |
A gang member failed for good. |
<k> of <n> workers failed; the first: … |
Failed |
Several failed. |
Worker <w>: Exceeded its time limit of <d> |
Failed |
A worker ran past time_limit_seconds. |
… ; worker group <g> cannot finish without it |
Failed |
A serving group failed for good in a complete_with run. |
All workers ended; some were cancelled |
Cancelled |
|
Quota reached: <why> |
any | A history entry (and an event) written the first time the run waits on its workspace's quota. |
Worker states#
| State | Meaning |
|---|---|
Pending |
Not running and waiting to be placed (or placed, and its machine has not picked it up yet). |
Preparing |
Its machine is setting it up: credentials, drives, network, identity. |
Pulling |
Its machine is pulling the image. |
ReadyToStart |
Gang runs: set up and waiting at the barrier for every member. |
Starting |
The container was started. |
Running |
Its machine reports it running and sends heartbeats. |
Stale |
Running, but its heartbeat is late. |
Stopping |
Asked to stop (cancelled, preempted, out of time); its machine is stopping the container. |
Completed |
Exited 0. Terminal unless restart_policy: Always. |
Failed |
Exited non-zero, could not start, or ran out of restarts. Terminal when it will not be restarted. |
Cancelled |
Stopped before or while running. Terminal, except after a preemption (it is queued again). |
Down |
Lost: its machine stopped answering, or its heartbeat expired. Restarted under its restart_policy. |
stateDiagram-v2
[*] --> Pending
Pending --> Preparing: placed, machine picks it up
Preparing --> Pulling
Preparing --> Starting
Pulling --> Starting
Preparing --> ReadyToStart: gang
Pulling --> ReadyToStart: gang
ReadyToStart --> Starting: every member ready
Starting --> Running
Running --> Stale: heartbeat late
Stale --> Running: heartbeat back
Stale --> Down: no heartbeat for its TTL
Running --> Down: machine lost
Running --> Completed: exit 0
Running --> Failed: exit ≠ 0, OOM
Preparing --> Failed: image, config, start errors
Pulling --> Failed
Running --> Stopping: stop, preemption, time limit
Stopping --> Cancelled
Pending --> Cancelled: stopped before placement
Failed --> Pending: restarted by policy
Down --> Pending: restarted by policy
Completed --> Pending: restart_policy Always
Cancelled --> Pending: preempted, queued again
Liveness#
Each running worker must send a heartbeat within heartbeat_ttl_seconds (default 60 s).
| From | Condition | To | Reason |
|---|---|---|---|
Running |
no heartbeat (after 15 s of grace) | Stale |
Heartbeat lease expired |
Stale |
heartbeat again, machine up | Running |
Heartbeat received |
Stale |
no heartbeat for its TTL | Down |
No heartbeat for the worker's full TTL after going stale |
| any active state | its machine is Down |
Down |
Machine lost: its lease expired |
| any active state | its machine was removed | Down |
Machine removed from the cluster |
Stopping |
its machine is Down or removed |
Cancelled |
Machine lost while stopping / Machine removed from the cluster while stopping |
Restarts#
A worker that ended is started again when its policy says so:
restart_policy |
Restarted after |
|---|---|
OnFailure (default) |
Failed or Down |
Always |
Failed, Down or Completed |
Never |
never (Down with Never fails for good) |
- Back-off: 10 s before the first restart, doubling each time, at most 5 min. While it waits, the worker shows when it restarts (
restart_at); the console shows Restarting in a loop. - Budget: 10 restarts. The next failure is final:
Retry budget exhausted after 10/10 restarts. - Reset: a worker that stays
Runningfor 10 min has its count reset to 0. - Who restarts: under
on_failure: RestartJob(the default for gangs) every worker is requeued:Gang restart: train-1 failed. UnderRestartTask, only the failed one:Restarted by policy (restart 2/10); in a gang, a leader failure still restarts all. - Not counted: a preempted worker is queued again without spending its budget:
Preempted by <run>; queued again.
Worker reasons from its machine#
| Reason | State | Meaning |
|---|---|---|
Waiting for the node network (subnet assignment and mesh) |
Preparing |
The machine's network is not ready yet. |
Waiting for its external port to be assigned |
Preparing |
An external access needs its port. |
Waiting for data volume <d>: <why> |
Preparing |
A drive is not mountable yet. |
Waiting for drive <d> to be filled on <machine> |
Preparing |
A placed drive's copy is being made. |
waiting for external secret <c> (key <k>) to be synced to this node |
Preparing |
A credential is not on the machine yet. |
waiting for registry secret <c> to be synced to this node |
Preparing |
The registry credential is not there yet. |
Waiting for the worker's identity: <why> |
Preparing |
identity: true and no certificate yet. |
Pulling <image> |
Pulling |
|
Waiting at the gang barrier |
ReadyToStart |
|
Container started |
Starting |
|
Container running |
Running |
|
Exited with code <n> |
Completed (0) or Failed |
|
OOMKilled: the container exceeded its memory limit |
Failed |
Raise memory_bytes (or per_gpu.memory_bytes). |
ErrImagePull: <error> |
Failed |
Wrong image name, tag, or registry credential. |
Image <image> is not present and the pull policy is None |
Failed |
|
Container create failed: <error> / Container start failed: <error> |
Failed |
The runtime refused the container. |
Container not found on the node |
Failed |
The container disappeared from the machine. |
Invalid task: <error> |
Failed |
A mount path or host path the machine refuses. |
Writing config <path>: <error> |
Failed |
|
Outbound policy: <why>, Agent sandbox: <why>, Model server: <why>, Interactive token: <why> |
Failed |
The feature could not be set up on the machine. |
Stopped on request |
Cancelled |
The machine confirmed a stop. |
Stop reasons and causes#
A stop moves a placed worker to Stopping (a Pending one straight to Cancelled); its machine sends the stop signal, waits for the grace period, kills it, and confirms Cancelled. Two stops carry a cause that decides what happens next:
| Cause | Reason | Then |
|---|---|---|
| Preemption | Preempted by higher-priority work (<run>) |
Queued again once gone (a gang once every member is gone); no restart counted. |
| Time limit | Exceeded its time limit of <d> |
Fails for good; never restarted. |
Other stop reasons: A gang member exhausted its restart budget, A worker failed, and the run fails on a failure, A gang member exceeded its time limit, A worker exceeded its time limit, Gang barrier timed out waiting for all members (requeued after 15 min at the barrier), Run failed / Run completed / Run cancelled, and a reason given to POST /tasks/<w>/stop (default Stopped by scheduler).
Pending reasons#
While a worker waits to be placed, its reason (and its run's) says why. The console turns these into a headline and advice; the table shows the exact text.
| Reason | Meaning | What to do |
|---|---|---|
No machine fits: <r> of <n> machines ready |
No machine passes every filter; placement.nodes lists each machine's reason. |
Read the per-machine reasons (how). |
Gang cannot be placed whole: <r> of <n> machines ready |
Not every gang member fits at once. | Wait, ask for fewer GPUs, or loosen topology rules. |
Cannot place the minimum to start together: … |
MinAvailable: the first min_available do not fit together. |
Lower min_available. |
… : no machines are registered |
The workspace can use no machine on this cluster. | Add a machine, or check the workspace's pools. |
Choose where to keep data on <machine> |
A placed drive needs a data location on that machine. | Choose one on the machine's page. |
Waits for namespace quota: gpus 6 in use + 4 asked > 8 |
Over the workspace's quota (Gang waits for … for a gang). Waiting on quota never blocks others. |
Wait for the workspace's other work to end, or ask an admin. |
Waits for more than the namespace's quota allows at all (…): raise the quota or ask for less |
The run alone exceeds the quota. | Ask for less, or raise the quota. |
Waits for worker group <g> to be running / healthy |
depends_on. |
|
Array runs at most <k> workers at once |
array.max_parallel. |
|
Waiting up to 6 min more for InfiniBand fabric ib-0 (could start now on Ethernet) |
topology.patience_seconds. |
|
Next to run (priority <p>): starts by <time> UTC as running work ends |
The head of the queue; room is held for it. | |
Next to run (priority <p>): waits for running work to end |
The head; running work has no time limit, so no start time is known. | |
Queued behind higher-priority work (<run>) |
A higher-priority run is the head. | |
Queued behind <run>, which starts by <time> UTC / Queued behind <run> |
The head has the same or lower priority and earlier claim; this run cannot fit around it. | Give the run a time limit so it can backfill. |
Preempting <k> lower-priority worker(s) of <runs> to start |
This run is the head and is stopping others. | |
Starting: waiting for <k> preempted worker(s) to stop |
Runs from another workspace are never named: such reasons say (a job in another namespace) instead.
Per-machine reasons#
placement.nodes[].reasons (the console's What keeps it out column) uses these:
| Reason | Meaning |
|---|---|
machine is not ready (<state>) / (no capacity reported) |
Down, or not yet reporting. |
under pressure: <condition> |
Memory or disk pressure: no new work. |
machine does not accept runs: <why> / machine is cordoned: <why> |
An admin stopped new work there. |
worker carries none of the labels this machine accepts |
The machine takes only work selecting its labels. |
machine is not among the selected names / machine lacks label <k>=<v> |
node_selection (pools included). |
no active RDMA port / no active InfiniBand port at 200 Gb/s or faster (…) |
network asked for RDMA. |
another worker of this run is already on the machine |
Workers spread one per machine. |
insufficient cpu: need 32 cores, 16 free |
|
insufficient memory: need <bytes> bytes, <bytes> free |
|
insufficient gpus: need 8, 4 free and usable |
|
needs 8 healthy GPUs, 6 free (2 more free but at risk) |
healthy_only. |
no H100 GPU, no nvidia H100 or H200 with at least 80 GB GPU |
No GPU on it matches gpu_requests. |
gpu fault on the machine: GPU 3: Xid 79… |
A GPU reported a fatal fault. |
the run is kept on one machine, <m> / the run is kept within rack <r> / no rack set on this machine (topology.astraeus.io/rack) / the run may use at most <n> machines / rack <r> already has a worker of this run |
topology rules. |
drive <d> does not exist / drive <d>: <why> / drive <d> is on <m>, in another site (<s>) / drive <d> needs <size> at this machine's data location, <size> free |
Drives. |
machine cannot enforce an outbound policy (…), machine cannot run sandboxed workers (…), machine serves no tool gateway to its workers (…) |
The run needs a capability the machine lacks. |
Exit codes#
| Exit | State | Usual cause |
|---|---|---|
0 |
Completed |
Success. |
1–125 |
Failed |
The program's own error. |
126, 127 |
Failed |
Command not executable, or not found in the image. |
137 (128 + 9) |
Failed |
Killed with SIGKILL. Out of memory is reported as OOMKilled instead. |
143 (128 + 15) |
Failed |
Exited on SIGTERM without handling it, outside a stop. |
A worker that Astraeus stops (cancel, preemption, time limit) ends Cancelled whatever its exit code; a time-limit stop then fails the run. The last failure's exit code, reason and last 50 lines of log (up to 16 KiB) are kept on the worker as last_failure.