Stalls and stragglers#
A worker can be Running and doing nothing useful: a collective that never completes, a deadlock, a data loader waiting on storage that does not answer. And a synchronous gang runs at the pace of its slowest worker: one GPU held down by heat slows every step of every machine. Astraeus watches for both, says which worker and why, and — when the run asks — restarts a stalled one. This page explains what is read, when a worker counts as stalled, and how to set it.
What counts as progress#
Each machine reads, for every worker running on it:
- Output: the tail of the worker's log changing (any line counts; nothing is read for meaning but NCCL's own words).
- Its GPUs: one of them 5% busy or more.
- NCCL: its watchdog's line in the log (
Watchdog caught collective operation timeout,NCCL communicator was aborted), and — for a worker of a multi-machine run that has gone quiet — NCCL's RAS port (2.24 and later, port 28028 inside the container) saying a communicator is incomplete or a rank is missing. - Its own metric, if the run names one (
resilience.stall.progress_metric): a value its metrics endpoint serves, such assteporsamples_total, changing.
When a worker is stalled#
| The run says | Stalled when |
|---|---|
| a progress metric | the metric has not changed for after_seconds (output and GPUs are not looked at). |
| nothing (GPU workers) | no output and its GPUs idle for after_seconds (default 1 800 s). |
| — and NCCL says a collective is stuck | no output for the shorter of after_seconds and 10 minutes since NCCL first said so. Busy GPUs are not progress then: a hung collective spins them at 100%. |
A worker whose machine says nothing of it (an agent from before this feature, a snapshot not in yet) is never judged stalled. Interactive runs (notebooks, environments), machines' checks and services (lifetime: Service) are not watched unless their run sets resilience.stall. Workers without GPUs are watched only when the run sets it.
When a worker is found stalled:
- It is marked: a history entry in the state it is in (
Running),Stalled: No progress for 12m: no output since 14:02 UTC; NCCL: NCCL RAS: rank 8 (node gpu-1, PID 4242) is missing: …, with the evidence indetails.stall. That entry is an event; your organisation's alert rules turn it into theworker_stalledalert. - With
action: restart, it is stopped — its stop signal, its grace,cause: stalledin its stop file — and restarted, a restart for a failure (ASTRAEUS_RESTART_CAUSE=stalled). In a gang that restarts together (on_failure: RestartJob), the gang restarts. - When progress resumes, it is said too (
Progress again: no longer stalled), and the alert resolves by the same key.
The GPU time a stalled worker held from when it stopped making progress is counted as stalled in its run's goodput — including the minutes before the stall was known, when its spinning GPUs looked busy.
Stragglers#
In a gang whose workers have all been running a while, Astraeus looks for one worker that stands out, strongest evidence first:
- Its own time: with
resilience.straggler.time_metric(a gauge of each worker's time per iteration on its metrics endpoint), the worker atratio(1.2) × the others' median or more. - A GPU slowed down: NVML says its clocks are held for heat, for power or by the hardware, or — among busy GPUs — its SM clock is under 85% of the gang's median:
GPU 3 slowed down by heat: 1386 MHz against 1980 MHz across the gang, 92 °C. - A slow link: its machine's slowest compute rail runs below what its rails are made for, or topology discovery found a degraded or miswired port on it.
- Weaker, said only beside a stronger sign or alone: a GPU at risk (heat, errors), its GPUs far busier than the others' (they wait for it).
A sign every worker shares names nobody. The same worker found on two passes in a row is named on the run: a history entry of the run (Straggler: Worker llm-2 (rank 2, on gpu-2) holds the gang back: …, details.straggler), an event, the run_straggler alert; and said over when the gang runs evenly again. Astraeus does not move the worker: that is a person's decision (drain the machine, or fence the GPU).
See it#
A run's page says a stalled worker and a straggler at the top, with the evidence and a link to the log; its Restarts tab shows each worker's last output and when its GPUs were last busy. The workspace's Goodput page lists every stalled worker and straggler now.
$ astra astraeus stalls
STALLED llm-1 (run llm, on gpu-1) since 14:02: no output since 14:02 UTC; NCCL: NCCL RAS: rank 8 (node gpu-1, PID 4242) is missing: collective #4242 (AllReduce) in progress on the others
STRAGGLER run sweep-7: Worker sweep-7-2 (rank 2, on gpu-6) holds the gang back: GPU 3 slowed down by heat: 1386 MHz against 1980 MHz across the gang, 92 °C
$ astra astraeus progress llm
$ curl -sS "$ASTRALYX_API/stalls" -H "Authorization: Bearer $ASTRALYX_TOKEN"
$ curl -sS "$ASTRALYX_API/runs/llm/progress" -H "Authorization: Bearer $ASTRALYX_TOKEN"
GET /stalls answers stalled (each with run, worker, machine and stall: since, detected_at, evidence, action) and stragglers (each with run, summary and straggler: worker, rank, machine, evidence, since).
Set it#
{
"stall": {"after_seconds": 900, "action": "restart"},
"straggler": {"time_metric": "step_seconds", "ratio": 1.25}
}
| Field | Type | Default | Description |
|---|---|---|---|
stall.enabled |
boolean | true |
false turns stall detection off for the run. |
stall.after_seconds |
integer, 30 to 604 800 | 1 800 | How long without progress. |
stall.action |
alert, restart |
alert |
alert: marked and said. restart: also stopped and restarted. |
stall.progress_metric |
metric name | — | A metric of the run's own whose change is progress; decides alone. |
straggler.time_metric |
metric name | — | A gauge of each worker's own time per iteration. |
straggler.ratio |
number, 1.01 to 10 | 1.2 | How far above the others' median. |
With the CLI: astra astraeus run … --stall-after 15m --restart-stalled.