Self-healing#
GPUs fail: a memory error, a GPU that falls off the bus, a driver that stops answering, a machine that loses power. Astralyx watches every machine for faults, keeps the faulty GPU (or the whole machine, when the fault is the machine's) out of new work at once, and repairs it step by step: it moves the work off, resets the GPU, reboots the machine, has a neighbour power-cycle it through its BMC, and — when nothing else works — collects signed evidence for its vendor. What disrupts nothing happens by itself; what would disrupt work waits for your approval, unless you let it act. Safety limits keep a burst of faults from taking a pool down, and every repair that ran leaves a signed receipt. Use this page to follow your machines' health, approve or reject repairs, drain machines, decide how far repairs go by themselves in each pool, connect a machine's BMC, and get RMA evidence.
Before you begin#
- Who. Reading health, incidents and repairs needs a workspace role
from
viewerup; a workspace sees the machines of its pools only. Deciding — approving or rejecting a repair, draining, maintenance, policies, BMCs, receipts and RMA evidence — needs the organisation roleowneroradmin, in the console, withastra, through the API, or through an AI assistant (which asks for your confirmation in the console first). - Nothing to install. The machine's agent reports faults the moment it
sees them and carries out repairs on its machine. A GPU reset uses
nvidia-smi(NVIDIA) oramd-smi(AMD), installed with the driver; a reboot usessystemctl reboot. A power cycle needs the machine's BMC to speak Redfish, and another machine that reaches it (see Power-cycle through the BMC). - Checks. A machine is checked before it returns to work after a repair, with the checks of Acceptance and checks: they run only while its GPUs are free, below every run.
- In the API examples,
ASTRALYX_APIishttps://api.astralyx.cloud/v1,ASTRALYX_TOKENan API token of an organisation admin for all their workspaces (the organisation's routes refuse a token scoped to one workspace), and<org>,<cluster>and<org>/<workspace>stand for yours. The examples show eight machines,gpu-01togpu-08, in the poolh100, each with eight H100 GPUs.
How it works#
flowchart LR
F[Fault seen] --> I[Incident]
I --> S{Next step}
S -->|acts by itself| R[Repair]
S -->|recommended or held| A[Waits for approval]
A -->|approved| R
R --> V[Checked]
V -->|healthy| OK[Back to work]
V -->|not fixed| S
S -->|nothing left| RMA[RMA evidence]
- A fault is seen. Each machine's agent reports a GPU's Xid (NVIDIA's error codes) as soon as the kernel logs it, and its GPUs' health, its disks' arrays, its network ports — and whether it reports at all.
- What is faulty is kept out of work at once. A GPU whose fault is its own — its Xid's action is to reset the GPU, or to restart the application — is fenced alone: the machine's other GPUs keep working. Only a fault that is the machine's — a GPU fallen off the bus (Xid 79), a GPU asking for the machine to be rebooted (Xid 154), a driver that does not answer — stops new work on the whole machine. A run that failed with the fault restarts once the fault is known, so it does not land on the faulty GPU again.
- An incident is opened, one per fault and what it is about (a GPU, a network port, the machine). The same fault seen again keeps the same incident open.
- The incident climbs a ladder of repairs (below), one step at a time. Each step runs at the level of autonomy its pool's policy gives it: done, proposed to you, or only said.
- Each repair is checked. A reset GPU must read healthy again; a rebooted or power-cycled machine must come back and pass its checks (the quick check, DCGM's diagnostic, GPU copy bandwidth and a short burn-in) before it takes work. A step that did not fix the fault is tried again (twice for a GPU reset and a reboot, once otherwise), then the next step is decided.
- The incident ends when a repair worked, or when the fault cleared by itself. When the ladder runs out, the machine waits for replacement, with its RMA evidence ready.
Every step that ran — by itself or approved — leaves a receipt, and every change of a machine's state, every incident and every repair is an event your alerts can send to e-mail, Slack, a webhook, PagerDuty or Opsgenie.
Machine states#
| State | New work | How it gets there |
|---|---|---|
| Ready | Yes | Joined, or back from a repair that passed its checks. |
| Suspect | No; what runs keeps running unless the fault ends it | A fault that needs more than one GPU fenced. |
| Draining | No; what runs finishes, or is moved off at the deadline | You drained it, or a repair needs it empty. |
| Quarantined | No | Out of service, waiting for a repair (yours, or an approved one). |
| Repairing | No | A GPU reset, reboot or power cycle is under way. |
| Validating | Only its checks | Checked before it returns to work. |
Beside these, a machine is Down (it stopped reporting), in Maintenance (an admin took it out of service), Accepting or Acceptance failed (see Acceptance and checks), or Idle (it never reported). Machine states has the rest.
Cordon, drain, maintenance. Cordoning only stops new work; what runs stays. Draining also empties the machine: what still runs at the deadline is stopped and queued again, without spending its restarts. Maintenance takes the machine out of service until you bring it back.
What is done about each fault#
| Fault | Seen as | Steps, in order |
|---|---|---|
Application error on a GPU (gpu-app) |
Xid 13, 31, 43, 94 | Restart the application: its run restarts it. The GPU is fine. |
GPU memory error (gpu-memory) |
Xid 48, 95; uncorrectable ECC errors | Fence the GPU, reset it (its bad page is retired), then RMA if it happens again. A row that could not be remapped (Xid 64): fence, then RMA. |
NVLink error (gpu-nvlink) |
Xid 74, 149; links down | Fence the GPU, reset it, drain, reboot, RMA. |
GPU firmware error (gpu-firmware) |
Xid 119, 120, and others whose action is a GPU reset | Fence the GPU, reset it, drain, reboot, power-cycle¹, RMA. |
GPU lost (gpu-lost) |
Xid 79 | Stop new work on the machine, drain, reboot, power-cycle¹, RMA. |
Driver not answering (driver) |
The GPUs cannot be read | Stop new work, drain, reboot, power-cycle¹, RMA. |
Machine not reporting (machine-down) |
Down for 2 minutes | Power-cycle¹, then RMA. Without a BMC: said only. |
Storage failed (storage) |
A RAID array or pool under its filesystems failed | Stop new work, drain, then a person rebuilds or reinstalls it. |
Network port in trouble (network) |
An RDMA port flapping, erroring or down | Network checks. Runs over RDMA rank the machine last meanwhile. |
GPU at risk (at-risk) |
Memory errors growing, PCIe replays, NVLink errors, a remap pending | Diagnostics (the deep suite), then fence the GPU, then RMA. |
GPU running hot (thermal) |
Slowed down by heat | Said, naming the straggler: a synchronous run goes at its pace. |
Check failed (check-failed) |
The quick check or a validation failed | Stop new work, diagnostics, RMA. |
Lemon (lemon) |
Fails far more often than its pool's peers | Said; large runs keep off it (see Lemons). |
Many faults at once (storm) |
Too many incidents in a pool within an hour | Automatic repairs pause in the pool (see Safety limits). |
¹ Only when the machine's BMC is set.
A GPU fault whose recovery action is the machine's reboot (Xid 154 saying so) takes the machine's steps instead: stop new work, drain, reboot, power-cycle¹, RMA — a GPU reset would not take.
On machines whose GPUs share an NVSwitch baseboard (HGX and DGX), a GPU is reset only with the machine drained: the driver resets the baseboard's GPUs together.
How far repairs go by themselves#
Each step of a fault's ladder runs at a level of autonomy:
| Level | What happens |
|---|---|
observe |
Only said: the incident and its recommendation. Nothing is proposed or done. |
recommend |
Proposed: a person approves it with one click. |
act-with-approval |
Prepared and held for approval; the organisation's admins are told. Unanswered, it expires after approval_wait_minutes (default 24 hours). |
act |
Done, within the pool's safety limits. |
A pool's policy sets a level per fault class (above) and a ceiling per step; a step gets the lower of the two. Without a policy, every class may act, and each step's ceiling is its default:
| Step | Default | Why |
|---|---|---|
Restart the application elsewhere (requeue) |
act |
Protects work; disrupts nothing the fault had not. |
Fence the GPU (fence-gpu) |
act |
|
Stop new work on the machine (mark-suspect) |
act |
What runs keeps running. |
Reset the GPU (gpu-reset) |
act |
Only an idle GPU is reset. |
Collect RMA evidence (rma) |
act |
Reads only. |
Run diagnostics (diagnose) |
recommend |
Costs GPU time. |
Drain the machine (drain) |
recommend |
Moves work off. |
Reboot (reboot) |
recommend |
The machine is drained first. |
Power-cycle through the BMC (power-cycle) |
recommend |
|
Reinstall (reimage) |
recommend |
Always a person's: approving it says it was done. |
Set a pool's policy#
- Open Organisation → Clusters & machines, the cluster, then Health.
- Under Policies, choose the pool (or the cluster's default,
*), and set each fault class's level, each step's ceiling and the limits. - Select Save.
classes:
gpu-lost: act
storage: observe
steps:
reboot: act
power-cycle: act-with-approval
max_actions_per_hour: 4
max_out_of_service_percent: 25
$ astra astraeus remediation-policy h100 --file h100.yaml
pool h100: at most 4 automatic action(s) an hour, 25% of the pool out of service by itself, 1 repair(s) at once; a storm is 10 incidents in an hour
gpu-lost: act
storage: observe
step power-cycle: at most act-with-approval
step reboot: at most act
astra astraeus remediation-policy shows every pool's;
--delete removes a pool's (the cluster's default then applies).
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-policies/h100" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"classes": {"gpu-lost": "act", "storage": "observe"}, "steps": {"reboot": "act", "power-cycle": "act-with-approval"}, "max_actions_per_hour": 4, "max_out_of_service_percent": 25}' \
| jq '{pool, classes, steps, max_actions_per_hour, max_out_of_service_percent, updated_by}'
{
"pool": "h100",
"classes": {"gpu-lost": "act", "storage": "observe"},
"steps": {"power-cycle": "act-with-approval", "reboot": "act"},
"max_actions_per_hour": 4,
"max_out_of_service_percent": 25,
"updated_by": "console:[email protected]"
}
GET …/remediation-policies lists every pool's (always with *, the
cluster's default); DELETE …/remediation-policies/h100 removes one.
A field the policy does not have, a class or step that is not one, or
reimage above recommend is refused (400 INVALID_REMEDIATION).
| Field | Default | Description |
|---|---|---|
classes |
every class act |
Level per fault class: observe, recommend, act-with-approval, act. |
steps |
the table above | The most each step gets, whatever the class. reimage at most recommend. |
max_actions_per_hour |
6 |
Automatic actions on machines or their work in the pool, an hour. |
max_out_of_service_percent |
10 |
Machines taken out of service by themselves, at most this share of the pool, and at least one machine. |
max_concurrent_repairs |
1 |
GPU resets, reboots and power cycles at once in the pool. |
storm_incidents_per_hour |
10 |
Incidents in the pool within an hour that pause automatic repairs. |
drain_deadline_minutes |
60 |
A drain a repair asks for moves what still runs off after this long. |
approval_wait_minutes |
1440 |
A step held for approval expires after this long (1 to 10080). |
retries |
gpu-reset: 2, reboot: 2, others 1 |
Attempts per step before the next step (1 to 10). |
lemon |
30 days, 3 incidents, 3×, 0.1 a day | When a machine is a lemon: window_days, min_incidents, factor, floor_per_day. |
Saving a policy is recorded in the organisation's audit log
(cluster.remediation.policy), naming you.
Safety limits#
A pool's limits keep repairs from doing more harm than the faults:
- Actions per hour. Past
max_actions_per_hour, the next automatic action waits for a person. - Out of service. Past
max_out_of_service_percentof the pool taken out by itself, the next waits for a person. A pool of a few machines has no spares to lose: at least one machine, never more than the share. - Repairs at once. Past
max_concurrent_repairs, a repair waits for the one under way. - Storms.
storm_incidents_per_hourincidents in a pool within an hour — a switch, a power feed, a bad driver release — trip its circuit breaker: automatic repairs pause in the pool, an incident of classstormis opened, and the alertremediation_pausedis sent. What would have acted waits for a person, saying why. Look for the cause the faults share, then resume.
A repair held by a limit says why (held: …) where it waits for approval.
Resume automatic repairs#
On the cluster's Health, the banner that says automatic repairs are paused in a pool has Resume automatic repairs.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-breakers/h100/reset" \
-H "Authorization: Bearer $ASTRALYX_TOKEN"
{"pool":"h100","reset":true,"by":"console:[email protected]"}
Only incidents opened after the reset count towards the next storm.
Follow the machines' health#
The cluster's Health (organisation admins), or Health in a workspace (its pools' machines): machines by state, those out of service and why, repairs waiting for approval, and incidents. Open an incident for its evidence, its steps and their states, and its timeline. A machine's page has its incident history.
$ astra astraeus health
machines: 1 Draining, 7 Ready
open incidents: 1; repairs waiting for a person: 1
MACHINE STATE SINCE INCIDENT WHY
gpu-03 Draining 2026-10-05 10:41:12 inc-20261005-92453955 GPU 3 of gpu-03: Xid 79 (GPU has fallen off the bus)
$ astra astraeus incidents inc-20261005-92453955
inc-20261005-92453955 on gpu-03 — Waiting (gpu-lost)
GPU 3 of gpu-03: Xid 79 — GPU has fallen off the bus (NVIDIA: reboot the machine)
recommended: GPU 3 fell off the bus. NVIDIA: reboot the machine; reseat the GPU, then replace it, if it happens again. Until the reboot none of its GPUs takes new work.
repairs:
act-20261005-84501b2c Succeeded mark-suspect Stop new work on the machine — GPU 3 on gpu-03 (by policy)
act-20261005-134f9e07 Succeeded drain Drain the machine — GPU 3 on gpu-03 (by console:[email protected])
act-20261005-5c2d8a11 Proposed reboot Reboot the machine — GPU 3 on gpu-03
astra astraeus health <machine> shows one machine with its
incidents; astra astraeus incidents lists the open ones (--all
for every one, --machine for one machine's). --json prints the
answers as the API gives them.
$ curl -sS "$ASTRALYX_API/machine-health" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>"
{"machines":{"Draining":1,"Ready":7},"out_of_service":[{"machine":"gpu-03","lifecycle":"Draining","reason":"GPU 3 of gpu-03: Xid 79 (GPU has fallen off the bus)","since":"2026-10-05T10:41:12Z","incident":"inc-20261005-92453955","pool":"h100"}],"open_incidents":1,"pending_actions":1}
GET /incidents (?state=all, ?machine=), GET /incidents/{id} (with
its repairs), GET /remediation-actions (those waiting for a person;
?state=open or ?state=all for more) and GET /machines/{name}/health
read the rest;
the organisation's routes
(/organizations/<org>/clusters/<cluster>/machine-health and the
same paths below it) read every machine of the organisation.
Approve or reject a repair#
A repair that waits for a person — Proposed, or AwaitingApproval —
says what it will do and why it waits: its level of autonomy, or the limit
that held it. The organisation's owners and admins are notified (in the
console, and by e-mail unless they turned it off), and the alert
remediation_approval is sent.
On the cluster's Health, under Waiting for approval, select Approve or Reject, with a reason if you like. A step that acts on the machine (a drain, a GPU reset, a reboot, a power cycle) says what will happen before you confirm.
$ astra astraeus repairs
REPAIR STATE STEP WHAT
act-20261005-5c2d8a11 Proposed reboot Reboot the machine — GPU 3 on gpu-03
$ astra astraeus repairs --approve act-20261005-5c2d8a11 --reason "drained, go"
act-20261005-5c2d8a11: Approved — Reboot the machine — GPU 3 on gpu-03
--reject <ID> rejects one: nothing is done, and its incident stays
open for a person.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-actions/act-20261005-5c2d8a11/approve" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"reason": "drained, go"}' | jq '{id, step, state, decided_by}'
{
"id": "act-20261005-5c2d8a11",
"step": "reboot",
"state": "Approved",
"decided_by": "console:[email protected]"
}
…/reject rejects it. Either answers 409 REMEDIATION_CONFLICT when
it no longer waits for a decision.
With an assistant connected (Connect AI assistants),
get_machine_health and list_incidents say what waits; the tool
approve_repair holds the approval for your confirmation in the
console and gives the assistant a link to show you. reject_repair
needs no confirmation: nothing is done to the machine.
An approved step runs as soon as it may: once the machine is drained, its
GPU idle, or a repair slot in the pool is free. Decisions are recorded in
the audit log (cluster.remediation.approve, cluster.remediation.reject)
and in the step's receipt.
An incident can also be closed (the fault is dealt with) or taken to its
next step (the step it waits on is passed over): from the incident in
the console, with astra astraeus incidents <id> --close or --escalate,
or with POST …/incidents/{id}/close and …/escalate.
Drain a machine#
A drain stops new work on the machine and empties it: what still runs at the deadline is stopped and queued again, without spending its restarts — a run that keeps checkpoints resumes from its last one. Without a deadline, what runs finishes. The machine stays drained until you undrain it.
In the cluster's Machines, select the machine's Drain, give a reason and a deadline (or let what runs finish), and confirm. Undrain puts it back to work.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/gpu-05/drain" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"reason": "PSU swap", "deadline_seconds": 1800}' | jq '{state, drain}'
{
"state": "Draining",
"drain": {"by": "console:[email protected]", "reason": "PSU swap", "since": "2026-10-05T11:02:40Z", "deadline": "2026-10-05T11:32:40Z"}
}
deadline_seconds is at most 30 days; 0 moves what runs off at once.
POST …/machines/gpu-05/undrain undrains it.
drain_machine holds the drain for your confirmation in the console;
undrain_machine puts the machine back to work.
Maintenance#
astra astraeus maintenance gpu-05 on --reason "BIOS update" (or
POST …/machines/gpu-05/maintenance) takes a machine out of service: it
takes no work, and is not repaired, until
astra astraeus maintenance gpu-05 off (DELETE …/machines/gpu-05/maintenance)
brings it back.
Power-cycle through the BMC#
A machine that stops answering cannot reboot itself. When its BMC is set,
Astralyx asks one of its neighbours — a machine of yours on the same
out-of-band network — to have the BMC power-cycle it (Redfish's
ComputerSystem.Reset: PowerCycle, or ForceRestart where the BMC does
not offer it). The BMC's username and password stay on your machines: you
keep them in a Credential with the keys
username and password, and only the neighbours named for the BMC
resolve it. The receipt names the neighbour that acted.
In the cluster's Machines, select the machine's BMC…, enter its Redfish address, the Credential, and the machines that reach it, then Save.
$ astra astraeus bmc gpu-05 --address https://10.0.0.21 --credential ops/bmc-admin --via gpu-04,gpu-06 --insecure-tls
astra astraeus bmc gpu-05 shows it; --remove forgets it.
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/gpu-05/bmc" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"address": "https://10.0.0.21", "credential": "ops/bmc-admin", "via": ["gpu-04", "gpu-06"], "insecure_tls": true}'
{"machine":"gpu-05","address":"https://10.0.0.21","credential":"ops/bmc-admin","via":["gpu-04","gpu-06"],"insecure_tls":true,"updated_by":"console:[email protected]","updated_at":"2026-10-05T11:10:02Z"}
| Field | Description |
|---|---|
address |
The BMC's Redfish root, https://<address>. |
system |
The Redfish system to reset (/redfish/v1/Systems/1); default: the first the BMC lists. |
credential |
<workspace>/<name>: a Credential of one of the organisation's workspaces, with the keys username and password. Its values are never sent to Astralyx: a request carrying username or password is refused (400 NO_SECRET_VALUES). |
via |
The machines that reach the BMC and may act for it: your organisation's, at least one — ideally two, in another rack. |
insecure_tls |
Accept the BMC's self-signed certificate (BMCs ship one; the address is on your out-of-band network). |
Power cycles are recommend by default. A pool that lets them act
(steps: {power-cycle: act}) gets a machine that lost power back without
anyone awake.
Receipts#
Every step that ran — by itself, under a policy, or approved — leaves a receipt: what it acted on and why (the fault's evidence, the machine's state before, the policy it was decided under), what was decided and by whom (the level of autonomy and why, a limit that held it, the person who approved and their reason), and how it ended (what the machine said, what was checked after, the neighbour that acted). Receipts are chained in order, each naming the previous one's digest, and signed by the cluster (a compact JWS, EdDSA): anyone with the cluster's public keys can check them offline.
The cluster's Health, Receipts: each receipt, with the checks the cluster makes of it.
$ astra astraeus repair-receipts
SEQ AT MACHINE STEP AUTONOMY OUTCOME BY
12 2026-10-05 10:58:31 gpu-03 reboot recommend Succeeded console:[email protected]
11 2026-10-05 10:44:02 gpu-03 drain recommend Succeeded console:[email protected]
10 2026-10-05 10:41:13 gpu-03 mark-suspect act Succeeded policy
$ astra astraeus repair-receipts --verify
10 ok
11 ok
12 ok
--verify checks each receipt's digest, its signature against the
cluster's keys (asked of the cluster, or --jwks <file>), and the
chain between the receipts listed, and exits non-zero if one fails.
GET /organizations/<org>/clusters/<cluster>/remediation-receipts
lists them, newest first (?limit=, default 50); …/remediation-receipts/{seq}
returns one with the cluster's checks of it. The cluster's public keys
are at GET /cluster-identity/jwks.
RMA evidence#
When a GPU or a machine is to be replaced, its incident's rma step
collects what its vendor asks for, as an evidence pack: the incident and its
timeline, the machine (model, driver, GPUs, its state when it failed), its
history, its Xids, its latest checks (DCGM's diagnostic, burn-in, copy
bandwidth), and NVIDIA's bug report from the machine itself when it can
still give one — with a manifest of every file's digest, signed by the
cluster. The organisation's admins are told (the alert machine_rma).
Open the incident on the cluster's Health and select Download RMA evidence.
GET /organizations/<org>/clusters/<cluster>/incidents/{id}/rma
returns the ZIP; 409 RMA_NOT_READY until the incident reached its
rma step.
Lemons#
A few machines fail far more often than their peers — about 1–2% of a fleet. A machine is a lemon when, over the policy's window (30 days), it had at least 3 hardware incidents and failed at 3 times its pool's median rate or more (and at least 0.1 a day). Its health says so, with a recommendation, and runs of 4 machines or more keep off it while they fit elsewhere: the more machines a run spans, the more one lemon costs it. A lemon still takes smaller work. Replace it, or have it serviced.
Alerts#
Self-healing adds these alerts:
| Alert | Sent when |
|---|---|
machine_incident |
An incident is opened, or resolved. |
machine_at_risk |
Something on a machine becomes at risk: a GPU (memory errors growing, heat), a network port (flapping, errors). |
remediation_approval |
A repair waits for approval. |
remediation_action |
A repair ran: done, or failed. |
machine_rma |
A machine waits for replacement: its RMA evidence is ready. |
remediation_paused |
Automatic repairs paused in a pool. |
Sent to PagerDuty or Opsgenie, an incident opens there and is resolved there when it is resolved here.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
A reboot stays Proposed |
Reboots are recommend by default. |
Approve it, or let the pool's policy act (steps: {reboot: act}). |
A repair waits with held: … |
A safety limit was reached. | Approve it, or raise the limit. |
| Automatic repairs paused | A storm tripped the pool's breaker. | Find the shared cause, then resume. |
| A machine that lost power stays Down | No BMC is set. | Set its BMC, or power it on. |
| A power cycle failed | The machines in via are down, or do not reach the BMC's address. |
Name machines on the out-of-band network, in another rack. |
| A GPU reset failed on an AMD machine | amd-smi is not installed. |
Install it with ROCm, or let the ladder reboot the machine. |
| A machine stays Validating | Its checks wait for its GPUs to be free, or failed. | See its validation in Acceptance and checks. |