Keep a fleet healthy over time#
A fleet does not stay healthy by itself between acceptance runs. This recipe sets a pool to check its idle machines on a schedule, so a degrading link or a GPU with growing memory errors turns up before a run finds it, shows you which machines fail more than their peers, and drains machines for a firmware upgrade without cancelling what is running there.
Before you begin#
- An organisation owner or admin: policies, drain and maintenance are theirs.
- A pool whose machines already passed acceptance once (Bring up a new GPU cluster).
- An organisation admin's API token,
curlandjq. Below,<org>and<cluster>stand for yours.
1. Check idle machines on a schedule#
Add periodic to the pool's validation policy: every so many hours,
Astraeus checks whichever of the pool's machines are idle at that moment —
never one with work on its GPUs, out of service, waiting for acceptance,
or in observe-only mode.
Compute → Validation → Policies, the pool h100, under
Periodic checks on idle machines, choose the suite (Quick or
Network) and Every N hours, Save.
$ curl -sS -X PUT "https://api.astralyx.cloud/v1/organizations/<org>/clusters/<cluster>/validation-policies/h100" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"periodic": {"suite": "quick", "every_hours": 12}}' | jq '{pool, periodic}'
{"pool": "h100", "periodic": {"suite": "quick", "every_hours": 12}}
quickruns each idle machine's quick check (driver, NVML, GPU count, memory, NVLink, PCIe, Xids, DCGM where installed) — seconds per machine, and it never blocks a run that needs the machine: checks always give way.networkinstead runs pairs and all-reduce among the machines that are idle together, catching a link that degraded since acceptance.- The longer suites (
acceptance,deep) are not run on a schedule — an admin starts those when they choose (Run acceptance). - A failed periodic check opens an incident, which flows into self-healing like any other fault.
Because a check never competes with work, a busy pool simply gets fewer of its machines checked in a given window; it catches up as machines go idle.
2. Find the fleet's lemons#
A few machines fail far more than their peers — about 1–2% of a fleet, in practice. Over a 30-day window, a machine with at least 3 hardware incidents and a failure rate at least 3 times its pool's median (and at least 0.1 a day) is a lemon. Astraeus marks it and keeps large runs off it while they fit elsewhere — the more machines a run spans, the more one lemon costs it — without taking it out of service on its own.
The cluster's Health lists lemons with their incident counts and a recommendation.
A lemon still takes smaller work; it is a signal to replace it or have it serviced, not an automatic quarantine. See Lemons.
3. Drain machines for a firmware upgrade#
Draining stops new work and empties the machine on your terms — what still runs at a deadline is stopped and queued again without spending its restart budget, rather than cancelled outright. A run that keeps checkpoints resumes from its last one.
On each machine, Drain, give a reason and a deadline (or let what runs finish), Confirm.
With no deadline_seconds (or let what runs finish in the console),
nothing is stopped — the machine simply takes no new work until it is
empty. deadline_seconds: 0 moves everything off at once.
Run the upgrade, then:
The machine's quick check runs within moments of coming back.
4. Take a machine fully out of service#
For longer work — a GPU swap, a rack move — maintenance keeps a machine out of service, and out of self-healing's repairs, until you say otherwise:
5. Do a rolling upgrade across a pool#
Drain, upgrade and undrain machines a few at a time, so the pool keeps most of its capacity throughout:
$ for m in gpu-01 gpu-02; do astra astraeus drain "$m" --deadline 30m --reason "firmware"; done
# … run the upgrade on gpu-01 and gpu-02 …
$ for m in gpu-01 gpu-02; do astra astraeus undrain "$m"; done
$ for m in gpu-03 gpu-04; do astra astraeus drain "$m" --deadline 30m --reason "firmware"; done
# …
Run a network check afterwards on the whole pool, to confirm the fabric came back the way it was:
What you get#
- Idle machines checked around the clock without anyone starting a validation by hand, and without ever delaying a run.
- The fleet's lemons named, so large runs avoid them and you know which machines to replace first.
- A firmware upgrade that moves through the pool a few machines at a time, with checkpointed work resuming rather than failing.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
| A machine's periodic check never seems to run | It is rarely idle, or is cordoned, in observe-only mode, or waiting for acceptance. | Nothing to fix: periodic checks skip busy or out-of-service machines by design. |
| A lemon still takes small runs | By design: lemons are not quarantined, only avoided by large runs when another machine fits. | Replace it or have it serviced if it keeps failing. |
| The reserved deadline passed and work is still running | The deadline only stops and requeues; it does not reserve time beyond it. | Give it a longer deadline, or stop the work yourself. |
astra astraeus undrain does nothing visible |
The machine was never drained, or is out of service for another reason (maintenance, acceptance). | Check its state on Compute → Machines. |