Self-healing for a production fleet#
A fleet of a few hundred GPUs faults constantly: a memory error here, a GPU that falls off the bus there, a machine that loses power overnight. You cannot have a person watch every one. This recipe sets a pool's policy so small faults repair themselves, larger ones page on-call with one-click approval, a machine that loses power comes back through its BMC without anyone awake, and every repair — automatic or approved — leaves a signed receipt you can verify months later for an audit or an RMA.
Before you begin#
- You are an owner or admin of the organisation: setting policies, BMCs, approving repairs and reading receipts are organisation decisions.
- Nothing to install on the machines: the agent already reports faults and carries out repairs. A power cycle needs the machine's BMC to speak Redfish, and at least one other machine (ideally two, in another rack) that reaches it on your out-of-band network.
- An organisation admin's API token,
curlandjq. Below,<org>and<cluster>stand for yours, andASTRALYX_APIishttps://api.astralyx.cloud/v1. The example: a poolh100of eight machines,gpu-01togpu-08.
1. Connect PagerDuty (or Opsgenie, or both)#
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/alerts/channels" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' \
-d '{"kind": "pagerduty", "name": "gpu-oncall", "key": "<Events API v2 routing key>", "region": "us"}'
{"id":"0192…","kind":"pagerduty","name":"gpu-oncall"}
Then a rule that sends self-healing's alerts, for every workspace and the machines, to that channel:
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/alerts/rules" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' \
-d '{"channel_id": "0192…", "workspace": null,
"events": ["machine_incident", "machine_at_risk", "remediation_approval", "remediation_action", "machine_rma", "remediation_paused"]}'
Each incident is keyed so the same fault is one PagerDuty incident, opened
and resolved from here; Opsgenie works the same way with
{"kind": "opsgenie", "key": "<API key>", "region": "eu"}. See
Alerts.
2. Connect each machine's BMC#
The BMC's credentials stay on your machines: put its username and password
in a Credential with the keys username
and password, then name the machines that may reach it.
On the machine's page, BMC…, enter its Redfish address, the
Credential, and the machines that reach it (gpu-02, gpu-03, from
another rack where you can), Save.
A request carrying username or password directly is refused
(400 NO_SECRET_VALUES): only the Credential's reference travels to
Astraeus. Repeat for every machine. See
Power-cycle through the BMC.
3. Set the pool's autonomy policy#
Decide, fault class by fault class and step by step, what the fleet may do by itself. This policy lets a fault's first steps — fencing a GPU, resetting it, stopping new work on the machine — happen automatically (their default), holds draining, rebooting and power-cycling for your approval, and caps how much the pool can do on its own:
classes:
gpu-lost: act
gpu-memory: act
storage: observe
steps:
power-cycle: act-with-approval
max_actions_per_hour: 4
max_out_of_service_percent: 25
max_concurrent_repairs: 1
storm_incidents_per_hour: 10
approval_wait_minutes: 60
$ astra astraeus remediation-policy h100 --file h100-policy.yaml
pool h100: at most 4 automatic action(s) an hour, 25% of the pool out of service by itself, 1 repair(s) at once; a storm is 10 incidents in an hour
gpu-lost: act
gpu-memory: act
storage: observe
step power-cycle: at most act-with-approval
max_out_of_service_percent: 25 and max_concurrent_repairs: 1 keep a
burst of faults from idling a quarter of the pool at once; past
storm_incidents_per_hour, the pool's circuit breaker trips and
automatic repairs pause until a person looks. See
How far repairs go by themselves
and Safety limits.
4. Let a fault run its course#
A GPU falls off the bus on gpu-03 (Xid 79). The machine stops taking new
work at once (mark-suspect, act by default) — what runs there keeps
running — and then waits for you: drain and reboot are both at their
default ceiling, recommend, so each is proposed and needs one approval:
$ astra astraeus health
open incidents: 1; repairs waiting for a person: 1
MACHINE STATE SINCE INCIDENT WHY
gpu-03 Suspect 2026-10-06 10:41:12 inc-20261006-92453955 GPU 3 of gpu-03: Xid 79 (GPU has fallen off the bus)
$ astra astraeus incidents inc-20261006-92453955
recommended: GPU 3 fell off the bus. NVIDIA: reboot the machine; reseat the GPU, then replace it, if it happens again. Until the reboot none of its GPUs takes new work.
repairs:
act-20261006-84501b2c Succeeded mark-suspect Stop new work on the machine — GPU 3 on gpu-03 (by policy)
act-20261006-134f9e07 Proposed drain Drain the machine — GPU 3 on gpu-03
PagerDuty opens an incident at the same moment, from
remediation_approval.
5. Approve the drain and reboot from your phone#
On the cluster's Health, under Waiting for approval, Approve (or Reject) each step, with a reason.
$ astra astraeus repairs --approve act-20261006-134f9e07 --reason "go"
act-20261006-134f9e07: Approved — Drain the machine — GPU 3 on gpu-03
$ astra astraeus repairs
REPAIR STATE STEP WHAT
act-20261006-5c2d8a11 Proposed reboot Reboot the machine — GPU 3 on gpu-03
$ astra astraeus repairs --approve act-20261006-5c2d8a11 --reason "drained, reboot it"
act-20261006-5c2d8a11: Approved — Reboot the machine — GPU 3 on gpu-03
The machine drains, reboots, and is checked (the quick check, DCGM's
diagnostic, copy bandwidth, a short burn-in) before it takes work again.
The PagerDuty incident resolves itself the moment this one does. To let
drain and reboot happen without a person — appropriate once you trust
the fleet's pattern of faults — raise their ceilings in the policy
(steps: {drain: act, reboot: act}).
6. When a machine loses power overnight#
gpu-05 stops answering at 2 a.m. With its BMC set and power-cycle:
act-with-approval, the step is prepared and held — but if you set
power-cycle: act instead, this runs with nobody awake: a neighbour
(gpu-04 or gpu-06) power-cycles it through Redfish
(ComputerSystem.Reset: PowerCycle), and the machine comes back and is
checked before it takes work again. The receipt names the neighbour that
acted.
7. When faults come in a storm#
A bad driver release, or a shared PSU, trips ten incidents in the pool
within an hour. The circuit breaker stops automatic repairs there, opens an
incident of class storm, and sends remediation_paused:
$ astra astraeus remediation-policy h100 --reset-breaker
pool h100: circuit breaker reset; automatic repairs resume
Find what the faults share before resetting it. See Safety limits.
8. Verify receipts, offline#
Every step that ran — automatic or approved — left a signed, chained receipt: what it acted on, why, what was decided and by whom, how it ended. Check them without trusting the console at all:
$ astra astraeus repair-receipts
SEQ AT MACHINE STEP AUTONOMY OUTCOME BY
12 2026-10-06 10:58:31 gpu-03 reboot recommend Succeeded console:[email protected]
11 2026-10-06 10:44:02 gpu-03 drain act Succeeded policy
10 2026-10-06 10:41:13 gpu-03 mark-suspect act Succeeded policy
$ astra astraeus repair-receipts --verify
10 ok
11 ok
12 ok
--verify checks each receipt's digest, its signature against the
cluster's public keys (fetched from GET /cluster-identity/jwks, or
--jwks <file> offline), and the chain between the receipts listed — the
same check an auditor runs months later with no access to the cluster at
all. See Receipts.
9. Get RMA evidence when a machine must be replaced#
$ astra astraeus incidents inc-20261006-92453955 --rma rma-gpu-03.zip
RMA evidence saved to rma-gpu-03.zip (2481920 bytes)
$ astra evidence verify rma-gpu-03.zip
The pack carries the incident's timeline, the machine's state, its Xids, its latest checks and NVIDIA's bug report where the machine could still give one, with a manifest of every file's digest, signed by the cluster — what most GPU vendors ask for with an RMA. See RMA evidence.
What you get#
- Fencing a faulty GPU and keeping new work off a suspect machine happen without paging anyone; draining, rebooting and power-cycling wait for one tap of Approve; a reimage always waits for a person.
- An incident in PagerDuty or Opsgenie the moment something needs a human, resolved automatically the moment it is fixed.
- A chain of signed receipts for every repair, and an RMA evidence pack ready the moment a machine is past saving — both verifiable with no access to Astraeus at all.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
A reboot stays Proposed forever |
Reboots are recommend by default. |
Approve it, or set steps: {reboot: act} if you want it automatic. |
A repair waits with held: … |
A safety limit was reached (actions per hour, out-of-service share, concurrent repairs). | Approve it, or raise the limit. |
| Automatic repairs paused in a pool | A storm tripped the circuit breaker. | Find the shared cause, then --reset-breaker. |
| A machine that lost power stays Down | No BMC is set for it. | Set its BMC, or power it on by hand. |
| A power cycle failed | The machines in via are themselves down, or do not reach the BMC's address. |
Name machines on the out-of-band network, ideally in another rack. |
astra evidence verify fails |
The ZIP was altered, or you used the wrong cluster's keys. | Re-download the pack; fetch --jwks from the right cluster. |