Skip to content

Self-healing#

GPUs fail: a memory error, a GPU that falls off the bus, a driver that stops answering, a machine that loses power. Astralyx watches every machine for faults, keeps the faulty GPU (or the whole machine, when the fault is the machine's) out of new work at once, and repairs it step by step: it moves the work off, resets the GPU, reboots the machine, has a neighbour power-cycle it through its BMC, and — when nothing else works — collects signed evidence for its vendor. What disrupts nothing happens by itself; what would disrupt work waits for your approval, unless you let it act. Safety limits keep a burst of faults from taking a pool down, and every repair that ran leaves a signed receipt. Use this page to follow your machines' health, approve or reject repairs, drain machines, decide how far repairs go by themselves in each pool, connect a machine's BMC, and get RMA evidence.

Before you begin#

  • Who. Reading health, incidents and repairs needs a workspace role from viewer up; a workspace sees the machines of its pools only. Deciding — approving or rejecting a repair, draining, maintenance, policies, BMCs, receipts and RMA evidence — needs the organisation role owner or admin, in the console, with astra, through the API, or through an AI assistant (which asks for your confirmation in the console first).
  • Nothing to install. The machine's agent reports faults the moment it sees them and carries out repairs on its machine. A GPU reset uses nvidia-smi (NVIDIA) or amd-smi (AMD), installed with the driver; a reboot uses systemctl reboot. A power cycle needs the machine's BMC to speak Redfish, and another machine that reaches it (see Power-cycle through the BMC).
  • Checks. A machine is checked before it returns to work after a repair, with the checks of Acceptance and checks: they run only while its GPUs are free, below every run.
  • In the API examples, ASTRALYX_API is https://api.astralyx.cloud/v1, ASTRALYX_TOKEN an API token of an organisation admin for all their workspaces (the organisation's routes refuse a token scoped to one workspace), and <org>, <cluster> and <org>/<workspace> stand for yours. The examples show eight machines, gpu-01 to gpu-08, in the pool h100, each with eight H100 GPUs.

How it works#

flowchart LR
  F[Fault seen] --> I[Incident]
  I --> S{Next step}
  S -->|acts by itself| R[Repair]
  S -->|recommended or held| A[Waits for approval]
  A -->|approved| R
  R --> V[Checked]
  V -->|healthy| OK[Back to work]
  V -->|not fixed| S
  S -->|nothing left| RMA[RMA evidence]
  1. A fault is seen. Each machine's agent reports a GPU's Xid (NVIDIA's error codes) as soon as the kernel logs it, and its GPUs' health, its disks' arrays, its network ports — and whether it reports at all.
  2. What is faulty is kept out of work at once. A GPU whose fault is its own — its Xid's action is to reset the GPU, or to restart the application — is fenced alone: the machine's other GPUs keep working. Only a fault that is the machine's — a GPU fallen off the bus (Xid 79), a GPU asking for the machine to be rebooted (Xid 154), a driver that does not answer — stops new work on the whole machine. A run that failed with the fault restarts once the fault is known, so it does not land on the faulty GPU again.
  3. An incident is opened, one per fault and what it is about (a GPU, a network port, the machine). The same fault seen again keeps the same incident open.
  4. The incident climbs a ladder of repairs (below), one step at a time. Each step runs at the level of autonomy its pool's policy gives it: done, proposed to you, or only said.
  5. Each repair is checked. A reset GPU must read healthy again; a rebooted or power-cycled machine must come back and pass its checks (the quick check, DCGM's diagnostic, GPU copy bandwidth and a short burn-in) before it takes work. A step that did not fix the fault is tried again (twice for a GPU reset and a reboot, once otherwise), then the next step is decided.
  6. The incident ends when a repair worked, or when the fault cleared by itself. When the ladder runs out, the machine waits for replacement, with its RMA evidence ready.

Every step that ran — by itself or approved — leaves a receipt, and every change of a machine's state, every incident and every repair is an event your alerts can send to e-mail, Slack, a webhook, PagerDuty or Opsgenie.

Machine states#

State New work How it gets there
Ready Yes Joined, or back from a repair that passed its checks.
Suspect No; what runs keeps running unless the fault ends it A fault that needs more than one GPU fenced.
Draining No; what runs finishes, or is moved off at the deadline You drained it, or a repair needs it empty.
Quarantined No Out of service, waiting for a repair (yours, or an approved one).
Repairing No A GPU reset, reboot or power cycle is under way.
Validating Only its checks Checked before it returns to work.

Beside these, a machine is Down (it stopped reporting), in Maintenance (an admin took it out of service), Accepting or Acceptance failed (see Acceptance and checks), or Idle (it never reported). Machine states has the rest.

Cordon, drain, maintenance. Cordoning only stops new work; what runs stays. Draining also empties the machine: what still runs at the deadline is stopped and queued again, without spending its restarts. Maintenance takes the machine out of service until you bring it back.

What is done about each fault#

Fault Seen as Steps, in order
Application error on a GPU (gpu-app) Xid 13, 31, 43, 94 Restart the application: its run restarts it. The GPU is fine.
GPU memory error (gpu-memory) Xid 48, 95; uncorrectable ECC errors Fence the GPU, reset it (its bad page is retired), then RMA if it happens again. A row that could not be remapped (Xid 64): fence, then RMA.
NVLink error (gpu-nvlink) Xid 74, 149; links down Fence the GPU, reset it, drain, reboot, RMA.
GPU firmware error (gpu-firmware) Xid 119, 120, and others whose action is a GPU reset Fence the GPU, reset it, drain, reboot, power-cycle¹, RMA.
GPU lost (gpu-lost) Xid 79 Stop new work on the machine, drain, reboot, power-cycle¹, RMA.
Driver not answering (driver) The GPUs cannot be read Stop new work, drain, reboot, power-cycle¹, RMA.
Machine not reporting (machine-down) Down for 2 minutes Power-cycle¹, then RMA. Without a BMC: said only.
Storage failed (storage) A RAID array or pool under its filesystems failed Stop new work, drain, then a person rebuilds or reinstalls it.
Network port in trouble (network) An RDMA port flapping, erroring or down Network checks. Runs over RDMA rank the machine last meanwhile.
GPU at risk (at-risk) Memory errors growing, PCIe replays, NVLink errors, a remap pending Diagnostics (the deep suite), then fence the GPU, then RMA.
GPU running hot (thermal) Slowed down by heat Said, naming the straggler: a synchronous run goes at its pace.
Check failed (check-failed) The quick check or a validation failed Stop new work, diagnostics, RMA.
Lemon (lemon) Fails far more often than its pool's peers Said; large runs keep off it (see Lemons).
Many faults at once (storm) Too many incidents in a pool within an hour Automatic repairs pause in the pool (see Safety limits).

¹ Only when the machine's BMC is set.

A GPU fault whose recovery action is the machine's reboot (Xid 154 saying so) takes the machine's steps instead: stop new work, drain, reboot, power-cycle¹, RMA — a GPU reset would not take.

On machines whose GPUs share an NVSwitch baseboard (HGX and DGX), a GPU is reset only with the machine drained: the driver resets the baseboard's GPUs together.

How far repairs go by themselves#

Each step of a fault's ladder runs at a level of autonomy:

Level What happens
observe Only said: the incident and its recommendation. Nothing is proposed or done.
recommend Proposed: a person approves it with one click.
act-with-approval Prepared and held for approval; the organisation's admins are told. Unanswered, it expires after approval_wait_minutes (default 24 hours).
act Done, within the pool's safety limits.

A pool's policy sets a level per fault class (above) and a ceiling per step; a step gets the lower of the two. Without a policy, every class may act, and each step's ceiling is its default:

Step Default Why
Restart the application elsewhere (requeue) act Protects work; disrupts nothing the fault had not.
Fence the GPU (fence-gpu) act
Stop new work on the machine (mark-suspect) act What runs keeps running.
Reset the GPU (gpu-reset) act Only an idle GPU is reset.
Collect RMA evidence (rma) act Reads only.
Run diagnostics (diagnose) recommend Costs GPU time.
Drain the machine (drain) recommend Moves work off.
Reboot (reboot) recommend The machine is drained first.
Power-cycle through the BMC (power-cycle) recommend
Reinstall (reimage) recommend Always a person's: approving it says it was done.

Set a pool's policy#

  1. Open Organisation → Clusters & machines, the cluster, then Health.
  2. Under Policies, choose the pool (or the cluster's default, *), and set each fault class's level, each step's ceiling and the limits.
  3. Select Save.
h100.yaml
classes:
  gpu-lost: act
  storage: observe
steps:
  reboot: act
  power-cycle: act-with-approval
max_actions_per_hour: 4
max_out_of_service_percent: 25
$ astra astraeus remediation-policy h100 --file h100.yaml
pool h100: at most 4 automatic action(s) an hour, 25% of the pool out of service by itself, 1 repair(s) at once; a storm is 10 incidents in an hour
  gpu-lost: act
  storage: observe
  step power-cycle: at most act-with-approval
  step reboot: at most act

astra astraeus remediation-policy shows every pool's; --delete removes a pool's (the cluster's default then applies).

$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-policies/h100" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"classes": {"gpu-lost": "act", "storage": "observe"}, "steps": {"reboot": "act", "power-cycle": "act-with-approval"}, "max_actions_per_hour": 4, "max_out_of_service_percent": 25}' \
  | jq '{pool, classes, steps, max_actions_per_hour, max_out_of_service_percent, updated_by}'
{
  "pool": "h100",
  "classes": {"gpu-lost": "act", "storage": "observe"},
  "steps": {"power-cycle": "act-with-approval", "reboot": "act"},
  "max_actions_per_hour": 4,
  "max_out_of_service_percent": 25,
  "updated_by": "console:[email protected]"
}

GET …/remediation-policies lists every pool's (always with *, the cluster's default); DELETE …/remediation-policies/h100 removes one. A field the policy does not have, a class or step that is not one, or reimage above recommend is refused (400 INVALID_REMEDIATION).

Field Default Description
classes every class act Level per fault class: observe, recommend, act-with-approval, act.
steps the table above The most each step gets, whatever the class. reimage at most recommend.
max_actions_per_hour 6 Automatic actions on machines or their work in the pool, an hour.
max_out_of_service_percent 10 Machines taken out of service by themselves, at most this share of the pool, and at least one machine.
max_concurrent_repairs 1 GPU resets, reboots and power cycles at once in the pool.
storm_incidents_per_hour 10 Incidents in the pool within an hour that pause automatic repairs.
drain_deadline_minutes 60 A drain a repair asks for moves what still runs off after this long.
approval_wait_minutes 1440 A step held for approval expires after this long (1 to 10080).
retries gpu-reset: 2, reboot: 2, others 1 Attempts per step before the next step (1 to 10).
lemon 30 days, 3 incidents, 3×, 0.1 a day When a machine is a lemon: window_days, min_incidents, factor, floor_per_day.

Saving a policy is recorded in the organisation's audit log (cluster.remediation.policy), naming you.

Safety limits#

A pool's limits keep repairs from doing more harm than the faults:

  • Actions per hour. Past max_actions_per_hour, the next automatic action waits for a person.
  • Out of service. Past max_out_of_service_percent of the pool taken out by itself, the next waits for a person. A pool of a few machines has no spares to lose: at least one machine, never more than the share.
  • Repairs at once. Past max_concurrent_repairs, a repair waits for the one under way.
  • Storms. storm_incidents_per_hour incidents in a pool within an hour — a switch, a power feed, a bad driver release — trip its circuit breaker: automatic repairs pause in the pool, an incident of class storm is opened, and the alert remediation_paused is sent. What would have acted waits for a person, saying why. Look for the cause the faults share, then resume.

A repair held by a limit says why (held: …) where it waits for approval.

Resume automatic repairs#

On the cluster's Health, the banner that says automatic repairs are paused in a pool has Resume automatic repairs.

$ astra astraeus remediation-policy h100 --reset-breaker
pool h100: circuit breaker reset; automatic repairs resume
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-breakers/h100/reset" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN"
{"pool":"h100","reset":true,"by":"console:[email protected]"}

Only incidents opened after the reset count towards the next storm.

Follow the machines' health#

The cluster's Health (organisation admins), or Health in a workspace (its pools' machines): machines by state, those out of service and why, repairs waiting for approval, and incidents. Open an incident for its evidence, its steps and their states, and its timeline. A machine's page has its incident history.

$ astra astraeus health
machines: 1 Draining, 7 Ready
open incidents: 1; repairs waiting for a person: 1

MACHINE                   STATE             SINCE                INCIDENT                  WHY
gpu-03                    Draining          2026-10-05 10:41:12  inc-20261005-92453955     GPU 3 of gpu-03: Xid 79 (GPU has fallen off the bus)
$ astra astraeus incidents inc-20261005-92453955
inc-20261005-92453955 on gpu-03 — Waiting (gpu-lost)
GPU 3 of gpu-03: Xid 79 — GPU has fallen off the bus (NVIDIA: reboot the machine)
recommended: GPU 3 fell off the bus. NVIDIA: reboot the machine; reseat the GPU, then replace it, if it happens again. Until the reboot none of its GPUs takes new work.

repairs:
  act-20261005-84501b2c         Succeeded         mark-suspect  Stop new work on the machine — GPU 3 on gpu-03 (by policy)
  act-20261005-134f9e07         Succeeded         drain         Drain the machine — GPU 3 on gpu-03 (by console:[email protected])
  act-20261005-5c2d8a11         Proposed          reboot        Reboot the machine — GPU 3 on gpu-03

astra astraeus health <machine> shows one machine with its incidents; astra astraeus incidents lists the open ones (--all for every one, --machine for one machine's). --json prints the answers as the API gives them.

$ curl -sS "$ASTRALYX_API/machine-health" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>"
{"machines":{"Draining":1,"Ready":7},"out_of_service":[{"machine":"gpu-03","lifecycle":"Draining","reason":"GPU 3 of gpu-03: Xid 79 (GPU has fallen off the bus)","since":"2026-10-05T10:41:12Z","incident":"inc-20261005-92453955","pool":"h100"}],"open_incidents":1,"pending_actions":1}

GET /incidents (?state=all, ?machine=), GET /incidents/{id} (with its repairs), GET /remediation-actions (those waiting for a person; ?state=open or ?state=all for more) and GET /machines/{name}/health read the rest; the organisation's routes (/organizations/<org>/clusters/<cluster>/machine-health and the same paths below it) read every machine of the organisation.

Approve or reject a repair#

A repair that waits for a person — Proposed, or AwaitingApproval — says what it will do and why it waits: its level of autonomy, or the limit that held it. The organisation's owners and admins are notified (in the console, and by e-mail unless they turned it off), and the alert remediation_approval is sent.

On the cluster's Health, under Waiting for approval, select Approve or Reject, with a reason if you like. A step that acts on the machine (a drain, a GPU reset, a reboot, a power cycle) says what will happen before you confirm.

$ astra astraeus repairs
REPAIR                        STATE             STEP          WHAT
act-20261005-5c2d8a11         Proposed          reboot        Reboot the machine — GPU 3 on gpu-03
$ astra astraeus repairs --approve act-20261005-5c2d8a11 --reason "drained, go"
act-20261005-5c2d8a11: Approved — Reboot the machine — GPU 3 on gpu-03

--reject <ID> rejects one: nothing is done, and its incident stays open for a person.

$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/remediation-actions/act-20261005-5c2d8a11/approve" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"reason": "drained, go"}' | jq '{id, step, state, decided_by}'
{
  "id": "act-20261005-5c2d8a11",
  "step": "reboot",
  "state": "Approved",
  "decided_by": "console:[email protected]"
}

…/reject rejects it. Either answers 409 REMEDIATION_CONFLICT when it no longer waits for a decision.

With an assistant connected (Connect AI assistants), get_machine_health and list_incidents say what waits; the tool approve_repair holds the approval for your confirmation in the console and gives the assistant a link to show you. reject_repair needs no confirmation: nothing is done to the machine.

An approved step runs as soon as it may: once the machine is drained, its GPU idle, or a repair slot in the pool is free. Decisions are recorded in the audit log (cluster.remediation.approve, cluster.remediation.reject) and in the step's receipt.

An incident can also be closed (the fault is dealt with) or taken to its next step (the step it waits on is passed over): from the incident in the console, with astra astraeus incidents <id> --close or --escalate, or with POST …/incidents/{id}/close and …/escalate.

Drain a machine#

A drain stops new work on the machine and empties it: what still runs at the deadline is stopped and queued again, without spending its restarts — a run that keeps checkpoints resumes from its last one. Without a deadline, what runs finishes. The machine stays drained until you undrain it.

In the cluster's Machines, select the machine's Drain, give a reason and a deadline (or let what runs finish), and confirm. Undrain puts it back to work.

$ astra astraeus drain gpu-05 --deadline 30m --reason "PSU swap"
gpu-05: Draining
$ astra astraeus undrain gpu-05
gpu-05: Ready
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/gpu-05/drain" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"reason": "PSU swap", "deadline_seconds": 1800}' | jq '{state, drain}'
{
  "state": "Draining",
  "drain": {"by": "console:[email protected]", "reason": "PSU swap", "since": "2026-10-05T11:02:40Z", "deadline": "2026-10-05T11:32:40Z"}
}

deadline_seconds is at most 30 days; 0 moves what runs off at once. POST …/machines/gpu-05/undrain undrains it.

drain_machine holds the drain for your confirmation in the console; undrain_machine puts the machine back to work.

Maintenance#

astra astraeus maintenance gpu-05 on --reason "BIOS update" (or POST …/machines/gpu-05/maintenance) takes a machine out of service: it takes no work, and is not repaired, until astra astraeus maintenance gpu-05 off (DELETE …/machines/gpu-05/maintenance) brings it back.

Power-cycle through the BMC#

A machine that stops answering cannot reboot itself. When its BMC is set, Astralyx asks one of its neighbours — a machine of yours on the same out-of-band network — to have the BMC power-cycle it (Redfish's ComputerSystem.Reset: PowerCycle, or ForceRestart where the BMC does not offer it). The BMC's username and password stay on your machines: you keep them in a Credential with the keys username and password, and only the neighbours named for the BMC resolve it. The receipt names the neighbour that acted.

In the cluster's Machines, select the machine's BMC…, enter its Redfish address, the Credential, and the machines that reach it, then Save.

$ astra astraeus bmc gpu-05 --address https://10.0.0.21 --credential ops/bmc-admin --via gpu-04,gpu-06 --insecure-tls

astra astraeus bmc gpu-05 shows it; --remove forgets it.

$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/gpu-05/bmc" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"address": "https://10.0.0.21", "credential": "ops/bmc-admin", "via": ["gpu-04", "gpu-06"], "insecure_tls": true}'
{"machine":"gpu-05","address":"https://10.0.0.21","credential":"ops/bmc-admin","via":["gpu-04","gpu-06"],"insecure_tls":true,"updated_by":"console:[email protected]","updated_at":"2026-10-05T11:10:02Z"}
Field Description
address The BMC's Redfish root, https://<address>.
system The Redfish system to reset (/redfish/v1/Systems/1); default: the first the BMC lists.
credential <workspace>/<name>: a Credential of one of the organisation's workspaces, with the keys username and password. Its values are never sent to Astralyx: a request carrying username or password is refused (400 NO_SECRET_VALUES).
via The machines that reach the BMC and may act for it: your organisation's, at least one — ideally two, in another rack.
insecure_tls Accept the BMC's self-signed certificate (BMCs ship one; the address is on your out-of-band network).

Power cycles are recommend by default. A pool that lets them act (steps: {power-cycle: act}) gets a machine that lost power back without anyone awake.

Receipts#

Every step that ran — by itself, under a policy, or approved — leaves a receipt: what it acted on and why (the fault's evidence, the machine's state before, the policy it was decided under), what was decided and by whom (the level of autonomy and why, a limit that held it, the person who approved and their reason), and how it ended (what the machine said, what was checked after, the neighbour that acted). Receipts are chained in order, each naming the previous one's digest, and signed by the cluster (a compact JWS, EdDSA): anyone with the cluster's public keys can check them offline.

The cluster's Health, Receipts: each receipt, with the checks the cluster makes of it.

$ astra astraeus repair-receipts
  SEQ  AT                   MACHINE               STEP          AUTONOMY            OUTCOME     BY
   12  2026-10-05 10:58:31  gpu-03                reboot        recommend           Succeeded   console:[email protected]
   11  2026-10-05 10:44:02  gpu-03                drain         recommend           Succeeded   console:[email protected]
   10  2026-10-05 10:41:13  gpu-03                mark-suspect  act                 Succeeded   policy
$ astra astraeus repair-receipts --verify
   10  ok
   11  ok
   12  ok

--verify checks each receipt's digest, its signature against the cluster's keys (asked of the cluster, or --jwks <file>), and the chain between the receipts listed, and exits non-zero if one fails.

GET /organizations/<org>/clusters/<cluster>/remediation-receipts lists them, newest first (?limit=, default 50); …/remediation-receipts/{seq} returns one with the cluster's checks of it. The cluster's public keys are at GET /cluster-identity/jwks.

RMA evidence#

When a GPU or a machine is to be replaced, its incident's rma step collects what its vendor asks for, as an evidence pack: the incident and its timeline, the machine (model, driver, GPUs, its state when it failed), its history, its Xids, its latest checks (DCGM's diagnostic, burn-in, copy bandwidth), and NVIDIA's bug report from the machine itself when it can still give one — with a manifest of every file's digest, signed by the cluster. The organisation's admins are told (the alert machine_rma).

Open the incident on the cluster's Health and select Download RMA evidence.

$ astra astraeus incidents inc-20261005-92453955 --rma rma-gpu-03.zip
RMA evidence saved to rma-gpu-03.zip (2481920 bytes); `astra evidence verify rma-gpu-03.zip` checks it
$ astra evidence verify rma-gpu-03.zip

GET /organizations/<org>/clusters/<cluster>/incidents/{id}/rma returns the ZIP; 409 RMA_NOT_READY until the incident reached its rma step.

Lemons#

A few machines fail far more often than their peers — about 1–2% of a fleet. A machine is a lemon when, over the policy's window (30 days), it had at least 3 hardware incidents and failed at 3 times its pool's median rate or more (and at least 0.1 a day). Its health says so, with a recommendation, and runs of 4 machines or more keep off it while they fit elsewhere: the more machines a run spans, the more one lemon costs it. A lemon still takes smaller work. Replace it, or have it serviced.

Alerts#

Self-healing adds these alerts:

Alert Sent when
machine_incident An incident is opened, or resolved.
machine_at_risk Something on a machine becomes at risk: a GPU (memory errors growing, heat), a network port (flapping, errors).
remediation_approval A repair waits for approval.
remediation_action A repair ran: done, or failed.
machine_rma A machine waits for replacement: its RMA evidence is ready.
remediation_paused Automatic repairs paused in a pool.

Sent to PagerDuty or Opsgenie, an incident opens there and is resolved there when it is resolved here.

Troubleshooting#

Symptom Cause Fix
A reboot stays Proposed Reboots are recommend by default. Approve it, or let the pool's policy act (steps: {reboot: act}).
A repair waits with held: … A safety limit was reached. Approve it, or raise the limit.
Automatic repairs paused A storm tripped the pool's breaker. Find the shared cause, then resume.
A machine that lost power stays Down No BMC is set. Set its BMC, or power it on.
A power cycle failed The machines in via are down, or do not reach the BMC's address. Name machines on the out-of-band network, in another rack.
A GPU reset failed on an AMD machine amd-smi is not installed. Install it with ROCm, or let the ladder reboot the machine.
A machine stays Validating Its checks wait for its GPUs to be free, or failed. See its validation in Acceptance and checks.