Skip to content

Self-healing for a production fleet#

A fleet of a few hundred GPUs faults constantly: a memory error here, a GPU that falls off the bus there, a machine that loses power overnight. You cannot have a person watch every one. This recipe sets a pool's policy so small faults repair themselves, larger ones page on-call with one-click approval, a machine that loses power comes back through its BMC without anyone awake, and every repair — automatic or approved — leaves a signed receipt you can verify months later for an audit or an RMA.

Before you begin#

  • You are an owner or admin of the organisation: setting policies, BMCs, approving repairs and reading receipts are organisation decisions.
  • Nothing to install on the machines: the agent already reports faults and carries out repairs. A power cycle needs the machine's BMC to speak Redfish, and at least one other machine (ideally two, in another rack) that reaches it on your out-of-band network.
  • An organisation admin's API token, curl and jq. Below, <org> and <cluster> stand for yours, and ASTRALYX_API is https://api.astralyx.cloud/v1. The example: a pool h100 of eight machines, gpu-01 to gpu-08.

1. Connect PagerDuty (or Opsgenie, or both)#

$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/alerts/channels" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' \
    -d '{"kind": "pagerduty", "name": "gpu-oncall", "key": "<Events API v2 routing key>", "region": "us"}'
{"id":"0192…","kind":"pagerduty","name":"gpu-oncall"}

Then a rule that sends self-healing's alerts, for every workspace and the machines, to that channel:

$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/alerts/rules" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' \
    -d '{"channel_id": "0192…", "workspace": null,
         "events": ["machine_incident", "machine_at_risk", "remediation_approval", "remediation_action", "machine_rma", "remediation_paused"]}'

Each incident is keyed so the same fault is one PagerDuty incident, opened and resolved from here; Opsgenie works the same way with {"kind": "opsgenie", "key": "<API key>", "region": "eu"}. See Alerts.

2. Connect each machine's BMC#

The BMC's credentials stay on your machines: put its username and password in a Credential with the keys username and password, then name the machines that may reach it.

On the machine's page, BMC…, enter its Redfish address, the Credential, and the machines that reach it (gpu-02, gpu-03, from another rack where you can), Save.

$ astra astraeus bmc gpu-05 --address https://10.0.0.21 --credential ops/bmc-admin --via gpu-04,gpu-06 --insecure-tls
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/gpu-05/bmc" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' \
    -d '{"address": "https://10.0.0.21", "credential": "ops/bmc-admin", "via": ["gpu-04", "gpu-06"], "insecure_tls": true}'

A request carrying username or password directly is refused (400 NO_SECRET_VALUES): only the Credential's reference travels to Astraeus. Repeat for every machine. See Power-cycle through the BMC.

3. Set the pool's autonomy policy#

Decide, fault class by fault class and step by step, what the fleet may do by itself. This policy lets a fault's first steps — fencing a GPU, resetting it, stopping new work on the machine — happen automatically (their default), holds draining, rebooting and power-cycling for your approval, and caps how much the pool can do on its own:

h100-policy.yaml
classes:
  gpu-lost: act
  gpu-memory: act
  storage: observe
steps:
  power-cycle: act-with-approval
max_actions_per_hour: 4
max_out_of_service_percent: 25
max_concurrent_repairs: 1
storm_incidents_per_hour: 10
approval_wait_minutes: 60
$ astra astraeus remediation-policy h100 --file h100-policy.yaml
pool h100: at most 4 automatic action(s) an hour, 25% of the pool out of service by itself, 1 repair(s) at once; a storm is 10 incidents in an hour
  gpu-lost: act
  gpu-memory: act
  storage: observe
  step power-cycle: at most act-with-approval
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-policies/h100" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' \
    --data-binary @h100-policy.json

max_out_of_service_percent: 25 and max_concurrent_repairs: 1 keep a burst of faults from idling a quarter of the pool at once; past storm_incidents_per_hour, the pool's circuit breaker trips and automatic repairs pause until a person looks. See How far repairs go by themselves and Safety limits.

4. Let a fault run its course#

A GPU falls off the bus on gpu-03 (Xid 79). The machine stops taking new work at once (mark-suspect, act by default) — what runs there keeps running — and then waits for you: drain and reboot are both at their default ceiling, recommend, so each is proposed and needs one approval:

$ astra astraeus health
open incidents: 1; repairs waiting for a person: 1

MACHINE  STATE     SINCE                INCIDENT               WHY
gpu-03   Suspect   2026-10-06 10:41:12  inc-20261006-92453955  GPU 3 of gpu-03: Xid 79 (GPU has fallen off the bus)
$ astra astraeus incidents inc-20261006-92453955
recommended: GPU 3 fell off the bus. NVIDIA: reboot the machine; reseat the GPU, then replace it, if it happens again. Until the reboot none of its GPUs takes new work.

repairs:
  act-20261006-84501b2c  Succeeded  mark-suspect  Stop new work on the machine — GPU 3 on gpu-03 (by policy)
  act-20261006-134f9e07  Proposed   drain         Drain the machine — GPU 3 on gpu-03

PagerDuty opens an incident at the same moment, from remediation_approval.

5. Approve the drain and reboot from your phone#

On the cluster's Health, under Waiting for approval, Approve (or Reject) each step, with a reason.

$ astra astraeus repairs --approve act-20261006-134f9e07 --reason "go"
act-20261006-134f9e07: Approved — Drain the machine — GPU 3 on gpu-03
$ astra astraeus repairs
REPAIR                 STATE     STEP    WHAT
act-20261006-5c2d8a11  Proposed  reboot  Reboot the machine — GPU 3 on gpu-03
$ astra astraeus repairs --approve act-20261006-5c2d8a11 --reason "drained, reboot it"
act-20261006-5c2d8a11: Approved — Reboot the machine — GPU 3 on gpu-03
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-actions/act-20261006-5c2d8a11/approve" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'Content-Type: application/json' -d '{"reason": "drained, reboot it"}'

The machine drains, reboots, and is checked (the quick check, DCGM's diagnostic, copy bandwidth, a short burn-in) before it takes work again. The PagerDuty incident resolves itself the moment this one does. To let drain and reboot happen without a person — appropriate once you trust the fleet's pattern of faults — raise their ceilings in the policy (steps: {drain: act, reboot: act}).

6. When a machine loses power overnight#

gpu-05 stops answering at 2 a.m. With its BMC set and power-cycle: act-with-approval, the step is prepared and held — but if you set power-cycle: act instead, this runs with nobody awake: a neighbour (gpu-04 or gpu-06) power-cycles it through Redfish (ComputerSystem.Reset: PowerCycle), and the machine comes back and is checked before it takes work again. The receipt names the neighbour that acted.

7. When faults come in a storm#

A bad driver release, or a shared PSU, trips ten incidents in the pool within an hour. The circuit breaker stops automatic repairs there, opens an incident of class storm, and sends remediation_paused:

$ astra astraeus remediation-policy h100 --reset-breaker
pool h100: circuit breaker reset; automatic repairs resume

Find what the faults share before resetting it. See Safety limits.

8. Verify receipts, offline#

Every step that ran — automatic or approved — left a signed, chained receipt: what it acted on, why, what was decided and by whom, how it ended. Check them without trusting the console at all:

$ astra astraeus repair-receipts
 SEQ  AT                   MACHINE  STEP          AUTONOMY    OUTCOME    BY
  12  2026-10-06 10:58:31  gpu-03   reboot        recommend   Succeeded  console:[email protected]
  11  2026-10-06 10:44:02  gpu-03   drain         act         Succeeded  policy
  10  2026-10-06 10:41:13  gpu-03   mark-suspect  act         Succeeded  policy
$ astra astraeus repair-receipts --verify
 10  ok
 11  ok
 12  ok

--verify checks each receipt's digest, its signature against the cluster's public keys (fetched from GET /cluster-identity/jwks, or --jwks <file> offline), and the chain between the receipts listed — the same check an auditor runs months later with no access to the cluster at all. See Receipts.

9. Get RMA evidence when a machine must be replaced#

$ astra astraeus incidents inc-20261006-92453955 --rma rma-gpu-03.zip
RMA evidence saved to rma-gpu-03.zip (2481920 bytes)
$ astra evidence verify rma-gpu-03.zip

The pack carries the incident's timeline, the machine's state, its Xids, its latest checks and NVIDIA's bug report where the machine could still give one, with a manifest of every file's digest, signed by the cluster — what most GPU vendors ask for with an RMA. See RMA evidence.

What you get#

  • Fencing a faulty GPU and keeping new work off a suspect machine happen without paging anyone; draining, rebooting and power-cycling wait for one tap of Approve; a reimage always waits for a person.
  • An incident in PagerDuty or Opsgenie the moment something needs a human, resolved automatically the moment it is fixed.
  • A chain of signed receipts for every repair, and an RMA evidence pack ready the moment a machine is past saving — both verifiable with no access to Astraeus at all.

Troubleshooting#

Symptom Cause Fix
A reboot stays Proposed forever Reboots are recommend by default. Approve it, or set steps: {reboot: act} if you want it automatic.
A repair waits with held: … A safety limit was reached (actions per hour, out-of-service share, concurrent repairs). Approve it, or raise the limit.
Automatic repairs paused in a pool A storm tripped the circuit breaker. Find the shared cause, then --reset-breaker.
A machine that lost power stays Down No BMC is set for it. Set its BMC, or power it on by hand.
A power cycle failed The machines in via are themselves down, or do not reach the BMC's address. Name machines on the out-of-band network, ideally in another rack.
astra evidence verify fails The ZIP was altered, or you used the wrong cluster's keys. Re-download the pack; fetch --jwks from the right cluster.