Skip to content

Restarts, signals and stops#

A worker can be stopped or lost at any time: preempted by higher-priority work, moved off a machine being drained, past its time limit, its machine lost, a GPU it held failing, or restarted with its gang. This page says exactly what Astraeus guarantees when that happens, how a run tells it how to be stopped and how often to be restarted, and how to read why a worker stopped. Astraeus does not know what your run computes: saving and reading back its work is your code's, on its drives; Astraeus makes sure the code is told, given time, and started again where its drives are.

What a restarted worker keeps#

  • Its identity. The same worker name (<run>-<rank>), the same run, the same rank and the same group rank; for a gang, the same world size and rendezvous variables (RANK, WORLD_SIZE, MASTER_ADDR…). Only where it runs changes.
  • Its drives, mounted again at the same paths on whichever machine it lands, read-only or read-write as before. When one cannot be mounted the worker waits in Preparing and its reason says why: Waiting for drive data: drive data does not exist (it was deleted, or never made), …: mounting it from gpu-3 failed: <why>, …: being mounted from gpu-3 (Exported), or Waiting for drive <d> to be filled on <machine> while a placed drive's copy is not whole yet.
  • Its specification: image, command, environment, credentials, requested resources.

It does not keep its container's own filesystem (anything written outside a drive), its processes, or its IP address.

How a worker is stopped#

  1. Its state becomes Stopping, with the reason (Preempted by higher-priority work (urgent), Moved off machine gpu-2: it is being drained, Exceeded its time limit of 2h, Stalled, restarted by its run's policy: …).
  2. The machine writes why to the stop file, $ASTRAEUS_STOP_FILE (/var/run/astraeus/lifecycle/stop), before anything else:

    /var/run/astraeus/lifecycle/stop
    {
      "cause": "preempted",
      "reason": "Preempted by higher-priority work (urgent)",
      "signal": "SIGTERM",
      "grace_seconds": 120,
      "deadline": "2026-10-07T14:32:10.512Z",
      "attempt": 2,
      "restarts": true
    }
    
    Field Description
    cause preempted, moved (a drain, a fenced GPU), time-limit, stalled, or requested (a person, or its run ending).
    reason The reason, in words.
    signal The signal about to be sent.
    grace_seconds, deadline How long it has before it is killed, and when that is.
    attempt The attempt it is on, from 1.
    restarts Whether it will be started again (preempted, moved, stalled and restarted) or this is the end.
  3. The machine sends the run's stop signal to the container's main process: SIGTERM unless the run says otherwise (resilience.stop_signal).

  4. After the grace period it sends SIGKILL. The grace is the run's stop_grace_seconds (1 to 3 600), at most the machine's ceiling (one hour by default, ASTRAEUS_MAX_STOP_GRACE); a run that does not say gets the machine's grace (30 s, ASTRAEUS_STOP_GRACE).
  5. The worker is Cancelled. Every step is a history entry of the worker (and an event), with its cause.

A machine that is lost sends no signal: nothing runs there to receive it. Save periodically too.

What the environment says#

Every worker gets, besides its usual variables:

Variable Value
ASTRAEUS_ATTEMPT The attempt it is on, from 1. It goes up each time the worker is placed again.
ASTRAEUS_RESTART_CAUSE On a restart, why the previous attempt ended: failure, machine-lost, gpu-fault, preempted, moved, stalled, gang (another worker of its gang failed), barrier (its gang's barrier timed out), requeued (by a person). Absent on attempt 1.
ASTRAEUS_RESTART_REASON The same, in words: Exited with code 1, Machine lost: unreachable for the whole grace period, Preempted by higher-priority work (urgent)…
ASTRAEUS_STOP_FILE /var/run/astraeus/lifecycle/stop (read-only; absent until a stop begins).

A small handler that saves on the signal and reads why:

handler.py
import json, os, signal, sys

def on_stop(signum, frame):
    why = {}
    try:
        with open(os.environ["ASTRAEUS_STOP_FILE"]) as f:
            why = json.load(f)
    except OSError:
        pass
    print(f"stopping ({why.get('cause', 'unknown')}): saving within {why.get('grace_seconds', 30)} s", flush=True)
    # save your state to a drive here, then exit
    sys.exit(0)

signal.signal(signal.SIGTERM, on_stop)
print(f"attempt {os.environ.get('ASTRAEUS_ATTEMPT', '1')}, "
      f"after {os.environ.get('ASTRAEUS_RESTART_CAUSE', 'nothing')}", flush=True)

Restarts are counted per cause#

Cause Spends Limit (run's field) Default
Its own failure: an exit other than 0, out of memory, its container could not start its restarts resilience.max_restarts 10
A stall its run restarts (Stalls and stragglers) its restarts resilience.max_restarts 10
Its machine lost or restarted under it, or a GPU it held failed (the GPU is fenced) its allowance for faults of machines resilience.max_machine_restarts 20
Preemption, a move off a drained machine nothing resilience.max_preemptions stops further preemption no cap
Its gang restarting for another worker, its gang's barrier timing out, a person's requeue nothing — —
  • A worker is restarted after a failure only when its restart_policy says so (OnFailure, the default; Always; never with Never). Preemption and moves restart it whatever its policy.
  • Back-off: 10 s before the first restart, doubling, at most 5 minutes (resilience.backoff_seconds, resilience.max_backoff_seconds). While it waits, the worker shows when it restarts.
  • Ten minutes running resets both counts of failures and of machine faults. Preemptions are never reset.
  • Spent: the worker fails for good with Retry budget exhausted after 10/10 restarts for its own failures; the last: Exited with code 1 (or …for faults of its machine or GPUs), and a gang stops the rest of its workers.
  • Minimum runtime: with resilience.preemptible_after_seconds, a run is not preempted before each attempt has run that long. The urgent run waits and says so: Next to run (priority 10): waits for running work to end; the lower-priority work it would preempt is protected (train: in its minimum runtime until 14:20 UTC).
  • Preemption cap: once preempted resilience.max_preemptions times, a run is not preempted again: … is protected (train: preempted as often as its run allows).

Each restart is a history entry of the worker with its cause and what it spent (details.restart: cause, attempt, budget, count), and a reason in words: Restarted after its machine was lost (restart 1 of 20), Gang restart after train-3: a GPU fault.

Set it on a run#

Write resilience and stop_grace_seconds in the run's specification (New run, then the specification). A run's page shows each worker's attempt, restarts by cause and how its last attempt ended on the Restarts tab.

$ astra astraeus run --name sft --image nvcr.io/nvidia/pytorch:25.01-py3 --gpus 8 \
    --restarts 5 --stop-signal SIGUSR1 --stop-grace 2m --min-runtime 30m \
    -- python train.py
$ astra astraeus progress sft
RANK   WORKER                     STATE        MACHINE            ATTEMPT  RESTARTS (F/M/P)        NOTE
0      sft-0                      Running      gpu-2                    2  0/1/0                   previous attempt: gpu-fault (GPU 3 on gpu-1: Xid 79 (GPU has fallen off the bus))
Flag Description
--restarts <n> Restart after its own failures up to n times (restart_policy: OnFailure).
--stop-signal <SIG> The stop signal.
--stop-grace <d> The grace period (30s, 2m).
--min-runtime <d> Not preempted before each attempt has run this long.
--stall-after <d>, --restart-stalled See Stalls and stragglers.
run.json
{
  "metadata": {"name": "sft"},
  "spec": {
    "task_template": {
      "image": "nvcr.io/nvidia/pytorch:25.01-py3",
      "command": "python",
      "args": ["train.py"],
      "restart_policy": "OnFailure",
      "stop_grace_seconds": 120,
      "resilience": {
        "stop_signal": "SIGUSR1",
        "max_restarts": 5,
        "max_machine_restarts": 20,
        "max_preemptions": 3,
        "preemptible_after_seconds": 1800,
        "backoff_seconds": 30,
        "max_backoff_seconds": 600
      },
      "requested_resources": {"gpu_requests": {"count": 8}, "cpu_cores": 32, "memory_bytes": 274877906944}
    }
  }
}
$ curl -sS -X POST "$ASTRALYX_API/runs" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Content-Type: application/json" -d @run.json
$ curl -sS "$ASTRALYX_API/runs/sft/progress" -H "Authorization: Bearer $ASTRALYX_TOKEN"

GET /runs/{name}/progress answers each worker's attempt, restarts (failures, machine, preemptions), previous (its cause, reason, machine, ended_at), stall and progress, and the run's straggler.

Reference#

Field (task_template.resilience) Type Default Description
stop_signal SIGTERM, SIGINT, SIGQUIT, SIGHUP, SIGUSR1, SIGUSR2 SIGTERM Sent first on a stop.
max_restarts integer, 0 to 1000 10 Restarts after its own failures and stalls.
max_machine_restarts integer, 0 to 1000 20 Restarts after faults of its machine or GPUs.
max_preemptions integer, 0 to 1000 no cap Preemptions (and moves) after which it is not preempted again.
preemptible_after_seconds integer, 0 to 604 800 none Not preempted before each attempt has run this long.
backoff_seconds integer, 1 to 86 400 10 The first wait before a restart, doubled each time.
max_backoff_seconds integer, 1 to 86 400, ≥ backoff_seconds 300 The longest wait.
stall object on for GPU workers See Stalls and stragglers.
straggler object — See Stalls and stragglers.

The cluster's defaults (10 and 20) are its operators' (ASTRAEUS_RETRY_MAX_ATTEMPTS, ASTRAEUS_MACHINE_RESTARTS on the scheduler).

Troubleshooting#

Symptom Cause Fix
Retry budget exhausted after 10/10 restarts for its own failures Your code fails each time. Read the last failure's output (the run's page, or last_failure); search the logs.
… for faults of its machine or GPUs Machines or GPUs keep failing under it. See the machines' self-healing; a machine failing far more than its peers is a lemon.
A worker waits in Preparing: Waiting for drive … Its drive is missing, its mount failed, or a placed copy is not whole. The reason says which; fix the drive, or wait for the copy.
The run does not save on a stop The signal did not reach your program. Start it with exec under a shell; or set stop_signal to the one your program handles.