Restarts, signals and stops#
A worker can be stopped or lost at any time: preempted by higher-priority work, moved off a machine being drained, past its time limit, its machine lost, a GPU it held failing, or restarted with its gang. This page says exactly what Astraeus guarantees when that happens, how a run tells it how to be stopped and how often to be restarted, and how to read why a worker stopped. Astraeus does not know what your run computes: saving and reading back its work is your code's, on its drives; Astraeus makes sure the code is told, given time, and started again where its drives are.
What a restarted worker keeps#
- Its identity. The same worker name (
<run>-<rank>), the same run, the same rank and the same group rank; for a gang, the same world size and rendezvous variables (RANK,WORLD_SIZE,MASTER_ADDR…). Only where it runs changes. - Its drives, mounted again at the same paths on whichever machine it lands, read-only or read-write as before. When one cannot be mounted the worker waits in
Preparingand its reason says why:Waiting for drive data: drive data does not exist (it was deleted, or never made),…: mounting it from gpu-3 failed: <why>,…: being mounted from gpu-3 (Exported), orWaiting for drive <d> to be filled on <machine>while a placed drive's copy is not whole yet. - Its specification: image, command, environment, credentials, requested resources.
It does not keep its container's own filesystem (anything written outside a drive), its processes, or its IP address.
How a worker is stopped#
- Its state becomes
Stopping, with the reason (Preempted by higher-priority work (urgent),Moved off machine gpu-2: it is being drained,Exceeded its time limit of 2h,Stalled, restarted by its run's policy: …). -
The machine writes why to the stop file,
$ASTRAEUS_STOP_FILE(/var/run/astraeus/lifecycle/stop), before anything else:/var/run/astraeus/lifecycle/stop{ "cause": "preempted", "reason": "Preempted by higher-priority work (urgent)", "signal": "SIGTERM", "grace_seconds": 120, "deadline": "2026-10-07T14:32:10.512Z", "attempt": 2, "restarts": true }Field Description causepreempted,moved(a drain, a fenced GPU),time-limit,stalled, orrequested(a person, or its run ending).reasonThe reason, in words. signalThe signal about to be sent. grace_seconds,deadlineHow long it has before it is killed, and when that is. attemptThe attempt it is on, from 1. restartsWhether it will be started again (preempted, moved, stalled and restarted) or this is the end. -
The machine sends the run's stop signal to the container's main process:
SIGTERMunless the run says otherwise (resilience.stop_signal). - After the grace period it sends
SIGKILL. The grace is the run'sstop_grace_seconds(1 to 3 600), at most the machine's ceiling (one hour by default,ASTRAEUS_MAX_STOP_GRACE); a run that does not say gets the machine's grace (30 s,ASTRAEUS_STOP_GRACE). - The worker is
Cancelled. Every step is a history entry of the worker (and an event), with its cause.
A machine that is lost sends no signal: nothing runs there to receive it. Save periodically too.
What the environment says#
Every worker gets, besides its usual variables:
| Variable | Value |
|---|---|
ASTRAEUS_ATTEMPT |
The attempt it is on, from 1. It goes up each time the worker is placed again. |
ASTRAEUS_RESTART_CAUSE |
On a restart, why the previous attempt ended: failure, machine-lost, gpu-fault, preempted, moved, stalled, gang (another worker of its gang failed), barrier (its gang's barrier timed out), requeued (by a person). Absent on attempt 1. |
ASTRAEUS_RESTART_REASON |
The same, in words: Exited with code 1, Machine lost: unreachable for the whole grace period, Preempted by higher-priority work (urgent)… |
ASTRAEUS_STOP_FILE |
/var/run/astraeus/lifecycle/stop (read-only; absent until a stop begins). |
A small handler that saves on the signal and reads why:
import json, os, signal, sys
def on_stop(signum, frame):
why = {}
try:
with open(os.environ["ASTRAEUS_STOP_FILE"]) as f:
why = json.load(f)
except OSError:
pass
print(f"stopping ({why.get('cause', 'unknown')}): saving within {why.get('grace_seconds', 30)} s", flush=True)
# save your state to a drive here, then exit
sys.exit(0)
signal.signal(signal.SIGTERM, on_stop)
print(f"attempt {os.environ.get('ASTRAEUS_ATTEMPT', '1')}, "
f"after {os.environ.get('ASTRAEUS_RESTART_CAUSE', 'nothing')}", flush=True)
Restarts are counted per cause#
| Cause | Spends | Limit (run's field) | Default |
|---|---|---|---|
| Its own failure: an exit other than 0, out of memory, its container could not start | its restarts | resilience.max_restarts |
10 |
| A stall its run restarts (Stalls and stragglers) | its restarts | resilience.max_restarts |
10 |
| Its machine lost or restarted under it, or a GPU it held failed (the GPU is fenced) | its allowance for faults of machines | resilience.max_machine_restarts |
20 |
| Preemption, a move off a drained machine | nothing | resilience.max_preemptions stops further preemption |
no cap |
| Its gang restarting for another worker, its gang's barrier timing out, a person's requeue | nothing | — | — |
- A worker is restarted after a failure only when its
restart_policysays so (OnFailure, the default;Always; never withNever). Preemption and moves restart it whatever its policy. - Back-off: 10 s before the first restart, doubling, at most 5 minutes (
resilience.backoff_seconds,resilience.max_backoff_seconds). While it waits, the worker shows when it restarts. - Ten minutes running resets both counts of failures and of machine faults. Preemptions are never reset.
- Spent: the worker fails for good with
Retry budget exhausted after 10/10 restarts for its own failures; the last: Exited with code 1(or…for faults of its machine or GPUs), and a gang stops the rest of its workers. - Minimum runtime: with
resilience.preemptible_after_seconds, a run is not preempted before each attempt has run that long. The urgent run waits and says so:Next to run (priority 10): waits for running work to end; the lower-priority work it would preempt is protected (train: in its minimum runtime until 14:20 UTC). - Preemption cap: once preempted
resilience.max_preemptionstimes, a run is not preempted again:… is protected (train: preempted as often as its run allows).
Each restart is a history entry of the worker with its cause and what it spent (details.restart: cause, attempt, budget, count), and a reason in words: Restarted after its machine was lost (restart 1 of 20), Gang restart after train-3: a GPU fault.
Set it on a run#
Write resilience and stop_grace_seconds in the run's specification (New run, then the specification). A run's page shows each worker's attempt, restarts by cause and how its last attempt ended on the Restarts tab.
$ astra astraeus run --name sft --image nvcr.io/nvidia/pytorch:25.01-py3 --gpus 8 \
--restarts 5 --stop-signal SIGUSR1 --stop-grace 2m --min-runtime 30m \
-- python train.py
$ astra astraeus progress sft
RANK WORKER STATE MACHINE ATTEMPT RESTARTS (F/M/P) NOTE
0 sft-0 Running gpu-2 2 0/1/0 previous attempt: gpu-fault (GPU 3 on gpu-1: Xid 79 (GPU has fallen off the bus))
| Flag | Description |
|---|---|
--restarts <n> |
Restart after its own failures up to n times (restart_policy: OnFailure). |
--stop-signal <SIG> |
The stop signal. |
--stop-grace <d> |
The grace period (30s, 2m). |
--min-runtime <d> |
Not preempted before each attempt has run this long. |
--stall-after <d>, --restart-stalled |
See Stalls and stragglers. |
{
"metadata": {"name": "sft"},
"spec": {
"task_template": {
"image": "nvcr.io/nvidia/pytorch:25.01-py3",
"command": "python",
"args": ["train.py"],
"restart_policy": "OnFailure",
"stop_grace_seconds": 120,
"resilience": {
"stop_signal": "SIGUSR1",
"max_restarts": 5,
"max_machine_restarts": 20,
"max_preemptions": 3,
"preemptible_after_seconds": 1800,
"backoff_seconds": 30,
"max_backoff_seconds": 600
},
"requested_resources": {"gpu_requests": {"count": 8}, "cpu_cores": 32, "memory_bytes": 274877906944}
}
}
}
$ curl -sS -X POST "$ASTRALYX_API/runs" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Content-Type: application/json" -d @run.json
$ curl -sS "$ASTRALYX_API/runs/sft/progress" -H "Authorization: Bearer $ASTRALYX_TOKEN"
GET /runs/{name}/progress answers each worker's attempt, restarts (failures, machine, preemptions), previous (its cause, reason, machine, ended_at), stall and progress, and the run's straggler.
Reference#
Field (task_template.resilience) |
Type | Default | Description |
|---|---|---|---|
stop_signal |
SIGTERM, SIGINT, SIGQUIT, SIGHUP, SIGUSR1, SIGUSR2 |
SIGTERM |
Sent first on a stop. |
max_restarts |
integer, 0 to 1000 | 10 | Restarts after its own failures and stalls. |
max_machine_restarts |
integer, 0 to 1000 | 20 | Restarts after faults of its machine or GPUs. |
max_preemptions |
integer, 0 to 1000 | no cap | Preemptions (and moves) after which it is not preempted again. |
preemptible_after_seconds |
integer, 0 to 604 800 | none | Not preempted before each attempt has run this long. |
backoff_seconds |
integer, 1 to 86 400 | 10 | The first wait before a restart, doubled each time. |
max_backoff_seconds |
integer, 1 to 86 400, ≥ backoff_seconds |
300 | The longest wait. |
stall |
object | on for GPU workers | See Stalls and stragglers. |
straggler |
object | — | See Stalls and stragglers. |
The cluster's defaults (10 and 20) are its operators' (ASTRAEUS_RETRY_MAX_ATTEMPTS, ASTRAEUS_MACHINE_RESTARTS on the scheduler).
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
Retry budget exhausted after 10/10 restarts for its own failures |
Your code fails each time. | Read the last failure's output (the run's page, or last_failure); search the logs. |
… for faults of its machine or GPUs |
Machines or GPUs keep failing under it. | See the machines' self-healing; a machine failing far more than its peers is a lemon. |
A worker waits in Preparing: Waiting for drive … |
Its drive is missing, its mount failed, or a placed copy is not whole. | The reason says which; fix the drive, or wait for the copy. |
| The run does not save on a stop | The signal did not reach your program. | Start it with exec under a shell; or set stop_signal to the one your program handles. |