Skip to content

Keep a fleet healthy over time#

A fleet does not stay healthy by itself between acceptance runs. This recipe sets a pool to check its idle machines on a schedule, so a degrading link or a GPU with growing memory errors turns up before a run finds it, shows you which machines fail more than their peers, and drains machines for a firmware upgrade without cancelling what is running there.

Before you begin#

  • An organisation owner or admin: policies, drain and maintenance are theirs.
  • A pool whose machines already passed acceptance once (Bring up a new GPU cluster).
  • An organisation admin's API token, curl and jq. Below, <org> and <cluster> stand for yours.

1. Check idle machines on a schedule#

Add periodic to the pool's validation policy: every so many hours, Astraeus checks whichever of the pool's machines are idle at that moment — never one with work on its GPUs, out of service, waiting for acceptance, or in observe-only mode.

Compute → Validation → Policies, the pool h100, under Periodic checks on idle machines, choose the suite (Quick or Network) and Every N hours, Save.

$ curl -sS -X PUT "https://api.astralyx.cloud/v1/organizations/<org>/clusters/<cluster>/validation-policies/h100" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"periodic": {"suite": "quick", "every_hours": 12}}' | jq '{pool, periodic}'
{"pool": "h100", "periodic": {"suite": "quick", "every_hours": 12}}
  • quick runs each idle machine's quick check (driver, NVML, GPU count, memory, NVLink, PCIe, Xids, DCGM where installed) — seconds per machine, and it never blocks a run that needs the machine: checks always give way.
  • network instead runs pairs and all-reduce among the machines that are idle together, catching a link that degraded since acceptance.
  • The longer suites (acceptance, deep) are not run on a schedule — an admin starts those when they choose (Run acceptance).
  • A failed periodic check opens an incident, which flows into self-healing like any other fault.

Because a check never competes with work, a busy pool simply gets fewer of its machines checked in a given window; it catches up as machines go idle.

2. Find the fleet's lemons#

A few machines fail far more than their peers — about 1–2% of a fleet, in practice. Over a 30-day window, a machine with at least 3 hardware incidents and a failure rate at least 3 times its pool's median (and at least 0.1 a day) is a lemon. Astraeus marks it and keeps large runs off it while they fit elsewhere — the more machines a run spans, the more one lemon costs it — without taking it out of service on its own.

The cluster's Health lists lemons with their incident counts and a recommendation.

$ astra astraeus health
machines: 1 Draining, 7 Ready
lemons (fail far more than their peers): gpu-06
$ astra astraeus health gpu-06
a lemon: 5 hardware faults in 30 days (0.42 a day, 4.2× its pool's median of 0.10): large runs keep off it; replace it or have it serviced

A lemon still takes smaller work; it is a signal to replace it or have it serviced, not an automatic quarantine. See Lemons.

3. Drain machines for a firmware upgrade#

Draining stops new work and empties the machine on your terms — what still runs at a deadline is stopped and queued again without spending its restart budget, rather than cancelled outright. A run that keeps checkpoints resumes from its last one.

On each machine, Drain, give a reason and a deadline (or let what runs finish), Confirm.

$ astra astraeus drain gpu-05 --deadline 30m --reason "firmware upgrade"
gpu-05: Draining
$ curl -sS -X POST "https://api.astralyx.cloud/v1/organizations/<org>/clusters/<cluster>/machines/gpu-05/drain" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"reason": "firmware upgrade", "deadline_seconds": 1800}'

With no deadline_seconds (or let what runs finish in the console), nothing is stopped — the machine simply takes no new work until it is empty. deadline_seconds: 0 moves everything off at once.

Run the upgrade, then:

$ astra astraeus undrain gpu-05
gpu-05: Ready

The machine's quick check runs within moments of coming back.

4. Take a machine fully out of service#

For longer work — a GPU swap, a rack move — maintenance keeps a machine out of service, and out of self-healing's repairs, until you say otherwise:

$ astra astraeus maintenance gpu-05 on --reason "GPU swap"
$ astra astraeus maintenance gpu-05 off

5. Do a rolling upgrade across a pool#

Drain, upgrade and undrain machines a few at a time, so the pool keeps most of its capacity throughout:

$ for m in gpu-01 gpu-02; do astra astraeus drain "$m" --deadline 30m --reason "firmware"; done
# … run the upgrade on gpu-01 and gpu-02 …
$ for m in gpu-01 gpu-02; do astra astraeus undrain "$m"; done
$ for m in gpu-03 gpu-04; do astra astraeus drain "$m" --deadline 30m --reason "firmware"; done
# …

Run a network check afterwards on the whole pool, to confirm the fabric came back the way it was:

$ astra astraeus validate --suite network --pool h100 --wait

What you get#

  • Idle machines checked around the clock without anyone starting a validation by hand, and without ever delaying a run.
  • The fleet's lemons named, so large runs avoid them and you know which machines to replace first.
  • A firmware upgrade that moves through the pool a few machines at a time, with checkpointed work resuming rather than failing.

Troubleshooting#

Symptom Cause Fix
A machine's periodic check never seems to run It is rarely idle, or is cordoned, in observe-only mode, or waiting for acceptance. Nothing to fix: periodic checks skip busy or out-of-service machines by design.
A lemon still takes small runs By design: lemons are not quarantined, only avoided by large runs when another machine fits. Replace it or have it serviced if it keeps failing.
The reserved deadline passed and work is still running The deadline only stops and requeues; it does not reserve time beyond it. Give it a longer deadline, or stop the work yourself.
astra astraeus undrain does nothing visible The machine was never drained, or is out of service for another reason (maintenance, acceptance). Check its state on Compute → Machines.