Acceptance and checks#
Astralyx checks your machines so that a broken GPU, a slow link or a failing disk shows before a run finds it. Each machine checks itself in seconds when it joins. Before new machines go to work, acceptance runs GPU diagnostics, a burn-in, copy bandwidth and the data disk on each machine, then measures the fabric between pairs of machines and an all-reduce across them; network checks measure the fabric alone, whenever you want. Every number is set against what the hardware should give, and the Cluster Report says what passed, what did not, and what to fix. Checks never compete with your work: they start only on free machines, below every run, and give way within seconds when work needs the GPUs. Use this page to run acceptance and network checks, read the report, and decide, pool by pool, whether new machines must pass before they take work.
Before you begin#
- Who. Reading checks, the report and the policies needs a workspace
role from
viewerup (auditorhas no access); a workspace sees the machines of its pools only. Starting and cancelling checks, setting policies, accepting a machine without its passing and converting one from observe-only mode need the organisation roleowneroradmin, in the console, withastra, or through the API. - Images. Checks run in containers on your machines, from two images
the machines pull:
ghcr.io/astralyx-cloud/astraeus-checks(nccl-tests, perftest, nvbandwidth, gpu-burn and fio; a-rocmvariant with rccl-tests for AMD GPUs) and NVIDIA's DCGM imagenvcr.io/nvidia/cloud-native/dcgm:4.6.1-1-ubuntu24.04, for the GPU diagnostic. A machine that cannot pull them cannot run those checks. - DCGM on the machine is optional. The quick check adds DCGM's shortest diagnostic where DCGM is installed; without it, it reads NVML alone. Acceptance needs nothing installed: its GPU diagnostic brings DCGM in its image.
- Between machines, pairs and all-reduce run on the machines' own
network, as multi-machine runs do: the machines must reach each other
over TCP on their addresses. A machine whose runs are on a bridged or
mesh network cannot be paired; its fabric checks come out
inconclusive, saying why. - A data location on a machine is where the disk check writes; a machine without one is not checked there (Drives and data).
- Topology places the fabric checks — pairs inside and across leaf groups, all-reduce inside one leaf group first — and gives what they are expected to measure: see Topology.
- In the API examples,
ASTRALYX_APIishttps://api.astralyx.cloud/v1,ASTRALYX_TOKENan API token of an organisation admin for all their workspaces (the organisation's routes refuse a token scoped to one workspace), and<org>,<cluster>and<org>/<workspace>stand for yours. The examples show the cluster of Topology: eight machines,gpu-01togpu-08in the poolh100, each with eight H100 GPUs and eight 400 Gb/s InfiniBand rails, in two leaf groups (gpu-01togpu-04,gpu-05togpu-08) under one spine group.
The quick check#
When a machine joins, and each time it comes back Up after being down (a reboot, an outage), Astralyx asks it to check itself. The agent answers from what it already reads, in seconds, without loading the machine:
| Item | Fails when | Warns when |
|---|---|---|
driver |
The GPU driver is not loaded, or does not answer within its deadline. | — |
nvml |
NVML calls block: the driver is wedged. | NVML (libnvidia-ml) cannot be loaded; the GPUs are read through nvidia-smi. |
gpu-count |
A GPU on the PCI bus is not one the driver answers for — fallen off the bus, or never bound: 7 of 8 GPUs on the PCI bus answer the driver: 0000:3a:00.0 (no driver). |
— |
memory |
Uncorrectable memory errors, or a failed row remap. | A row remap is pending: reset the GPU when idle. |
nvlink |
NVLink links down (GPU 4: 18 of 18 links down). |
— |
pcie |
— | A GPU's PCIe link below its width or speed (GPU 5: Gen5 x8 (of Gen5 x16)). |
xid |
A fatal fault the driver reported, latched by the agent (GPU 1 (Xid 79)). |
— |
dcgm |
DCGM's dcgmi diag -r 1 reports a failed test. |
It reports a warning. |
dcgmruns on NVIDIA machines wheredcgmiis on thePATHand DCGM's host engine (nv-hostengine) runs, for at most 90 seconds. Elsewhere it is skipped:DCGM is not installed: judged from NVML alone.- The verdict is the worst item's:
fail,warnorpass. It shows on the machine as its Quick check badge and in the report, with a fix for each item. Awarnor afailis also an event of the machine:Quick check: <what is wrong>. - It never blocks. A machine whose quick check fails still takes work; a GPU fault it reports fences that GPU as any fault does (GPUs).
- It is not made by itself on a machine in observe-only mode.
- A quick check that could not be made (the machine has not read itself yet) is tried again 2 minutes later. Astralyx waits at most 150 seconds for an answer.
Suites#
| Suite | What it runs | How long |
|---|---|---|
quick |
The quick check, now. | Seconds |
acceptance |
On each machine: the GPU diagnostic (DCGM, level 2), a 10-minute burn-in, GPU copy bandwidth and the data disk. Between machines: bandwidth and latency of pairs, rail by rail, then all-reduce over growing groups. | 30 to 60 minutes for a few machines: about 25 for two, 32 for eight, 38 for sixteen |
network |
The fabric only: pairs and all-reduce. | Minutes |
deep |
Acceptance with DCGM's level 3 and a 30-minute burn-in. | An hour or more |
A validation — one suite on some machines — runs in phases, each once the one before is over:
- The checks on each machine: one at a time on a machine, machines in parallel (up to the policy's limit).
- Pairs of machines: a pair starts as soon as both its machines are free, each machine in one pair at a time.
- The pairs that came out bad, again, with fewer at once.
- All-reduce, from two machines up, inside one leaf group before across the spines.
A machine that stays busy holds the next phase until its checks have run or are skipped. Name the machines or the pool to check, to leave busy machines out.
On each machine#
| Check | Name | Tool | Measures | Judged |
|---|---|---|---|---|
| GPU diagnostic (DCGM) | dcgm-diag |
dcgmi diag -r <level>, in NVIDIA's DCGM image |
Each test DCGM runs at that level | A test that fails is a fail, one that warns a warn. DCGM that could not run is inconclusive, not a verdict on the GPUs. Level 1 takes about a minute, 2 about 3 minutes, 3 about 30, 4 about an hour. |
| Burn-in | gpu-burn |
gpu-burn: every GPU at full load | Each GPU's errors, throughput (Gflop/s) and temperature | Wrong results: fail. Throughput under 90% of the median of the machine's GPUs: warn; under 80%: fail — a straggler, and every run it is in goes at its pace. At or above the model's slowdown temperature (87 °C for H100, H200, B200; 85 °C for A100, L40S, L4; 83 °C for GeForce RTX 4090): warn. |
| GPU copy bandwidth | nvbandwidth |
nvbandwidth | Host→GPU copies, GPU by GPU; GPU↔GPU copies over NVLink | Host→GPU against 87% of the GPU's PCIe link at its best (Gen5 x16: 54.8 GB/s); GPU↔GPU against 82% of its NVLink bandwidth (H100: 369 GB/s). Under 90% of that: fail. |
| Data disk | fio |
fio, one sequential job each way, in a directory of its own under the machine's data location (.astraeus-check-<id>), removed afterwards |
Read and write, GB/s | Under a floor for its medium: warn (NVMe 1.0 read, 0.5 write; SSD 0.3, 0.2; hard disk 0.1, 0.08; unknown 0.2, 0.1 GB/s). I/O errors: fail. |
The GPU diagnostic, the burn-in and copy bandwidth run on NVIDIA GPUs only. On a machine with AMD GPUs, acceptance checks the data disk, the pairs and all-reduce (with RCCL).
Between machines#
| Check | Name | Tool | Measures | Judged |
|---|---|---|---|---|
| Machine-to-machine bandwidth | ib-write-bw |
perftest's ib_write_bw and ib_write_lat, rail by rail, a few seconds each |
Each rail's write bandwidth (Gb/s); the pair's latency (µs) | Each rail against 97.5% of the slower end's rated rate (400 Gb/s: 390 Gb/s); between leaf groups, times 0.85 and divided by the leaf groups' oversubscription (1:1: 332 Gb/s), as for all-reduce. Under 90% of that: fail. Then pairs against each other (below). The other machine never answering on any rail: inconclusive. |
| All-reduce | nccl-allreduce |
nccl-tests' all_reduce (rccl-tests on AMD) |
Bus bandwidth (GB/s), and whether every value came back right | Against what the topology should give those machines (the expected all-reduce bandwidth): under 90% of it, fail. Wrong values, or NCCL finding no connection over the fabric: fail. Machines that could not start the test together: inconclusive. |
Pairs follow ClusterKit's method:
- Which pairs. Every pair of the machines checked: n − 1 rounds of
n/2 pairs, each machine in one pair a round. Eight machines are 28
pairs in 7 rounds. Past 64 machines, each is paired with 3 of the
others, its own leaf group's and another's in turn (
every_pairasks for all of them,peersfor more or fewer). Machines never tested before are also paired with 3 machines of the cluster that were, as a yardstick: a machine added to a large cluster is tested in minutes. A single machine checked where none was tested yet is paired with 3 of the others. - Bad pairs. Pairs are compared with their like — pairs inside one leaf group with each other, pairs across the spines with each other. A pair is bad when its slowest rail is under 93% of the best such pair's, or its latency over 2.1 times the lowest such pair's, or it failed its own test. A pair bad only against the others — every rail met its own expected value — is set on the path between the two machines, not on either of them.
- Retests. Bad pairs run again with half as many pairs at once, then half again, down to one at a time. A pair bad alone is bad; one that recovers was crowded by the others. A pair's last retest is its result.
Only machines with RDMA ports are paired, and only within a fabric: pairs never span fabrics that do not reach each other.
All-reduce runs over machines with GPUs: in each leaf group, groups of 2, 4, 8… machines and the whole group (when its size is not a power of two); then every machine together, across the spines, when they span several leaf groups. Machines whose leaf group is not known are a group of their own, judged as crossing the spines.
Verdicts and expected values#
| Verdict | Meaning |
|---|---|
pass |
As expected. |
warn |
Works, with something to look at: a GPU running hot, a slow disk. Counts as passing. |
fail |
Below what the hardware should give, or broken. |
inconclusive |
Not judged: the check could not run, or other work loaded the links it crossed. Never turned into a fail. |
A machine's verdict is its worst check's (fail, then inconclusive,
warn, pass), or untested when nothing ran. A pair's or a group's
fail counts against a machine only where a problem names it; an
inconclusive one counts against each of its machines. A pair bad only
against the other pairs fails the validation and is marked in the pairs'
matrix, but fails neither machine: the fault is on the path between
them. Each problem names
its subject — a machine, a GPU (gpu-05 GPU 2), a port, a pair — says
what is wrong, and gives a fix. Where topology
found the cause (a cable on another rail's leaf, a port below its rate, a
PCIe link below its width), the problem quotes the finding and its fix.
What a check should measure comes from the hardware itself, at the rates its links are rated for: a link that trained down does not lower what is expected of it.
| Measured | Expected | Passes at |
|---|---|---|
| Host→GPU copies | 87% of the GPU's PCIe link at its best | 90% of expected |
| GPU↔GPU copies | 82% of the GPU's NVLink bandwidth | 90% of expected |
| A rail between two machines | 97.5% of the slower end's rated rate | 90% of expected |
| All-reduce bus bandwidth | Rails × the slowest rail's line rate (Gb/s ÷ 8) × 0.92 within one leaf group; × 0.85 more through the spines, divided by the leaves' oversubscription | 90% of expected |
| Burn-in throughput | The median of the machine's GPUs | 90% (warn below), 80% (fail below) |
| Data disk | The floor for its medium | The floor (warn below) |
For eight 400 Gb/s rails, all-reduce is expected at 368 GB/s inside one leaf group and 313 GB/s across the spines of a non-blocking fabric. For a GPU model Astralyx does not know, copies are held to the PCIe and NVLink rates the machine reports, and its slowdown temperature is taken as 87 °C.
How checks stay out of your work's way#
A check runs as a run of the cluster's own, not of any workspace: it is not counted in any workspace's usage or quota. It asks for its machines whole — every GPU — so nothing else lands there while it runs, and:
- It starts only on free machines. A machine is checked only when no run holds any of its GPUs, it serves no inference replicas, nothing outside Astraeus uses one of its GPUs, and it hosts no deployment scaled to zero (a check could delay its wake-up; a policy can include such machines). It must be Up, not cordoned, not under pressure, and not in observe-only mode.
- One check at a time on a machine, and a few machines at once: at most
the policy's
max_concurrent_machinesof a pool's machines are under test together (8 by default). - It comes after every run. A check run's priority is below every run's. It is never the head of the queue, never takes room held for another run, and never preempts anything.
- It gives way within seconds. A run that needs its GPUs preempts it as it would any lower-priority work (Preemption); the check has 5 seconds to stop. Its step shows Preempted and runs again when the machine is free, a minute later at the soonest, so the work that needed it goes first; preempted 3 times, it is skipped.
- It never waits in the queue. A check run that could not be placed within 90 seconds (its machines got busy meanwhile) is withdrawn, and its step waits for them again. A step waits at most 6 hours for its machines, counted from when its phase opened and they were not free (and again after each preemption); then it is skipped.
- An explicit start (
explicit: true, by an organisation admin) may also check machines in observe-only mode, cordoned machines, and machines where something outside Astraeus was seen using a GPU. It still never takes a GPU another program holds, a machine serving inference replicas, or a GPU a run holds.
Fabric tests and other work#
- They run between free machines only.
- With other work on the cluster, only inside one leaf group: no spine is crossed, so no link another run uses is loaded.
- Across the spines — pairs under different leaf groups, all-reduce over
several — only in one of the policy's windows, when the
fabric is nearly idle (other work holds at most 5% of its GPUs and no run
spans machines), or when the request says
cross_spine: true. Until then the step waits:crosses the spines: waits for the policy's window, or the fabric nearly idle. - A result over links other work loaded — a run spanning leaf groups of the
same fabric while the test crossed the spines — is
inconclusive, neverfail:… — not judged: other work loaded the links it crossed (other work on 4 machines across leaf-r0-7-g0, leaf-r0-7-g1). - Each test lasts seconds per rail. The fabric is loaded for minutes only
when a request asks for a burn of the links (
fabric_burn: true): the pairs and all-reduce then run ten times longer. A policy never does.
When acceptance is offered#
Acceptance is offered, not imposed. Astralyx tells data-centre machines from workstations:
| Class | Machines |
|---|---|
datacenter |
Two GPUs or more, with RDMA ports or data-centre GPUs; not under WSL, not a Mac. |
workstation |
Every other machine: a single GPU, consumer GPUs without RDMA, WSL, a Mac. |
When at least two data-centre machines are Up, not in observe-only mode, and were never accepted, the Cluster Report offers acceptance:
Run acceptance before putting these machines to work? (~32 min for 8 machines, recommended for new clusters: GPU diagnostics, a burn-in, the fabric between them, all-reduce)
astra astraeus report prints the offer with the command that runs it:
$ astra astraeus report
8/8 machines pass.
Run acceptance before putting these machines to work? (~32 min for 8 machines, recommended for new clusters: GPU diagnostics, a burn-in, the fabric between them, all-reduce)
astra astraeus validate --suite acceptance --machines gpu-01,gpu-02,gpu-03,gpu-04,gpu-05,gpu-06,gpu-07,gpu-08
MACHINE STATE VERDICT QUICK DCGM BURN COPIES DISK PAIRS ALL-REDUCE
gpu-01 Ready pass pass — — — — — —
gpu-02 Ready pass pass — — — — — —
gpu-03 Ready pass pass — — — — — —
gpu-04 Ready pass pass — — — — — —
gpu-05 Ready pass pass — — — — — —
gpu-06 Ready pass pass — — — — — —
gpu-07 Ready pass pass — — — — — —
gpu-08 Ready pass pass — — — — — —
Workstations are never offered it, and no machine waits for acceptance
unless its pool's policy says so. You can
run it when the report offers it, or any time (below);
at install, with --accept; or for every new machine of a pool, with a
policy.
At install: --accept#
Add --accept to a machine's install command: acceptance starts by itself
once the machine joins, as any check, when its GPUs are free. Machines that
join together are checked together: acceptance starts 2 minutes after the
last of them joined. It runs once per machine, and does not keep the
machine from work. Where the machine's pool requires acceptance, the
policy's acceptance runs instead.
$ echo '9c4e…' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --control-plane https://connect.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme Research' --org-id 0192f0c4-7d1e-7a51-9c33-5e8b2a4f6d10 --name gpu-09 --accept && rm -f ./astraeus-token
See the installer reference.
Run acceptance#
- Open Compute → Validation. When the report offers acceptance, select Run acceptance in the offer; otherwise select Run checks.
- Choose the suite Acceptance, and the machines: by name, or a pool.
- Select Start. The validation's page follows each step: waiting, and why; running; done, with its verdict.
$ astra astraeus validate --suite acceptance --pool h100
validation val-20261005-7f3a2c: acceptance on gpu-01, gpu-02, gpu-03, gpu-04, gpu-05, gpu-06, gpu-07, gpu-08
https://console.astralyx.cloud/o/acme/w/vision/validation?cluster=lab-a&validation=val-20261005-7f3a2c
val-20261005-7f3a2c
The validation's id goes to standard output (with --json, the whole
validation), the rest to standard error. Without --machines or
--pool, every machine of the organisation on the cluster that is
Up is checked (in observe-only mode only with --explicit).
--wait follows it until it is over — each step's change on
standard error, then the validation with its steps — and exits
non-zero unless it Passed. Every flag is in the
CLI reference.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"suite": "acceptance", "pool": "h100", "note": "new rack r12"}' \
| jq '{id, suite, machines, reason, state, message}'
{
"id": "val-20261005-7f3a2c",
"suite": "acceptance",
"machines": [
"gpu-01",
"gpu-02",
"gpu-03",
"gpu-04",
"gpu-05",
"gpu-06",
"gpu-07",
"gpu-08"
],
"reason": "on-demand",
"state": "Pending",
"message": "Asked for (on-demand)"
}
The answer is 201, the whole validation, with requested_by naming
you (console:<e-mail>). Without machines, every machine of the
organisation on the cluster — of pool, when given — that is Up is
checked, those in observe-only mode only with explicit; machines
left out are named in message. Every field is in
Reference.
Each start is recorded in the organisation's audit log
(cluster.validation.start).
Follow it#
Compute → Validation lists recent validations. Select one: its steps, phase by phase, each with its machines, its state, why it waits, and its result.
$ astra astraeus validations
ID SUITE STATE MACHINES PASSED WHY AGE
val-20261005-7f3a2c acceptance Running 8 27/65 on-demand 39m
astra astraeus validations val-20261005-7f3a2c shows one, step by
step: each check, its machines, its state, its verdict and its result,
or why it waits; then the problems found, with their fixes.
$ curl -sS "$ASTRALYX_API/validations/val-20261005-7f3a2c" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>" \
| jq '{state, message, counts}'
{
"state": "Running",
"message": "Planned: 65 check run(s) on 8 machine(s)",
"counts": {
"steps": 65,
"waiting": 33,
"running": 3,
"passed": 27,
"warned": 0,
"failed": 1,
"inconclusive": 0,
"preempted": 1,
"skipped": 0
}
}
$ curl -sS "$ASTRALYX_API/validations/val-20261005-7f3a2c" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>" \
| jq -c '.steps[] | select(.state != "Done") | {id, state, why}' | head -5
{"id":"gpu-burn.gpu-03","state":"Preempted","why":"stopped for real work on gpu-03 at 10:41:07: it runs again when they are free"}
{"id":"gpu-burn.gpu-06","state":"Running","why":null}
{"id":"gpu-burn.gpu-07","state":"Running","why":null}
{"id":"nvbandwidth.gpu-08","state":"Running","why":null}
{"id":"ib-write-bw.gpu-01.gpu-08","state":"Waiting","why":null}
GET /validations lists them, newest first, without their steps.
| Validation state | Meaning |
|---|---|
Pending |
Asked for, not planned yet. |
Running |
Planned into steps; its checks run. |
Passed |
Every check passed (some may have warned). |
Failed |
A check failed. |
Inconclusive |
Nothing failed, but something could not be judged or run. |
Cancelled |
Stopped by a person. |
| Step state | Meaning |
|---|---|
Waiting |
For its phase, or for its machines to be free; why says which. |
Running |
Its check run is placed and running. |
Done |
Over: its result holds the verdict, the numbers against what was expected, the problems with fixes, and the end of the tool's own output. |
Preempted |
Real work needed its machines: stopped within seconds, it runs again when they are free. |
Skipped |
Not run: preempted 3 times, or its machines were not free in time. Counts as inconclusive. |
Cancelled |
The validation was cancelled. |
Cancel it#
Open the validation and select Cancel.
Its running checks stop and its open steps end Cancelled. A machine that
waited for that acceptance under a pool's policy is then Acceptance
failed (acceptance was cancelled: run it again, or accept the machine).
Cancelling is recorded in the audit log (cluster.validation.cancel).
Run network checks#
Run the network suite after recabling or replacing a switch, or when an all-reduce is slower than its run expected (Topology). It runs only the pairs and all-reduce.
Open Compute → Validation, select Run checks, choose the suite Network and the machines, and select Start. Test across the spines now runs those tests even while other work runs.
$ astra astraeus validate --suite network --machines gpu-02,gpu-03,gpu-04 --every-pair --wait
validation val-20261005-c41e09: network on gpu-02, gpu-03, gpu-04
https://console.astralyx.cloud/o/acme/w/vision/validation?cluster=lab-a&validation=val-20261005-c41e09
ib-write-bw gpu-03+gpu-04 Running
ib-write-bw gpu-03+gpu-04 Done pass: gpu-03 ↔ gpu-04: 8 rail(s), slowest 391 Gb/s, latency 1.63 µs
ib-write-bw gpu-02+gpu-04 Running
ib-write-bw gpu-02+gpu-04 Done pass: gpu-02 ↔ gpu-04: 8 rail(s), slowest 390 Gb/s, latency 1.64 µs
…
val-20261005-c41e09 network Passed (on-demand)
3/3 machines pass
CHECK MACHINES STATE VERDICT RESULT
ib-write-bw gpu-03+gpu-04 Done pass gpu-03 ↔ gpu-04: 8 rail(s), slowest 391 Gb/s, latency 1.63 µs
ib-write-bw gpu-02+gpu-04 Done pass gpu-02 ↔ gpu-04: 8 rail(s), slowest 390 Gb/s, latency 1.64 µs
ib-write-bw gpu-02+gpu-03 Done pass gpu-02 ↔ gpu-03: 8 rail(s), slowest 392 Gb/s, latency 1.62 µs
nccl-allreduce gpu-02+gpu-03 Done pass 2 machines: busbw 361.5 GB/s of 368 expected (98%; one leaf switch; expect ~368 GB/s all-reduce)
nccl-allreduce gpu-02+gpu-03+gpu-04 Done pass 3 machines: busbw 359.2 GB/s of 368 expected (98%; one leaf switch; expect ~368 GB/s all-reduce)
--cross-spine runs the tests across the spines now.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"suite": "network", "machines": ["gpu-02", "gpu-03", "gpu-04"], "options": {"every_pair": true}}' \
| jq -c '{id, state, machines}'
{"id":"val-20261005-c41e09","state":"Pending","machines":["gpu-02","gpu-03","gpu-04"]}
- On a cluster tested before, each machine named is paired with 3 peers
and with the other machines named;
peerssets how many peers, andevery_pair: truetests every pair of the machines named. cross_spine: trueruns the tests across the spines now. A result over links other work loaded meanwhile isinconclusive.fabric_burn: true(through the API) loads the links for minutes: the pairs and all-reduce run ten times longer. Only when you ask; a policy never does.- With the policy's
fabric_testsoff, or on machines without RDMA ports or GPUs, there is nothing to run: the validation endsInconclusive(Nothing to check: …).
Read the Cluster Report#
The Cluster Report puts every check's latest result together:
- A summary in a sentence: how many machines pass, what failed and its fix, the worst pair, and all-reduce over the most machines against what was expected.
- Counts: machines that pass, warn, fail, could not be judged, were not checked yet, wait to pass acceptance, or are in observe-only mode.
- The offer to run acceptance, when it applies (above).
- Machines × checks: each machine's state (
Ready,Accepting,AcceptanceFailed,Observe-only, orDown), its verdict, each check's verdict, and the topology's findings about it. - What to fix: each problem, where it is, and its fix.
- Measured against expected: every number beside the value it is held to.
- Pairs: the matrix of every pair tested — its slowest rail and its share of the best pair — with bad pairs marked.
- All-reduce: bus bandwidth by number of machines, measured against expected, inside one leaf group and across the spines.
Open Compute → Validation. The report is the page: the summary, the offer when there is one, the machines and their checks, what to fix, the pairs' matrix and the all-reduce curve. Download saves it as a printable page (HTML) or as JSON.
$ astra astraeus report
7/8 machines pass; gpu-05 GPU 2: 76% of its peers' throughput under load (a straggler: every run it is in goes at its pace) — Check its cooling and power (clocks held down), its PCIe link, and its NVLink; pair gpu-03–gpu-06 at 92% of the best pair; check the cables, transceivers and switch ports of both machines on this path; all-reduce over 8 machines at 93% of expected (291 of 313 GB/s).
MACHINE STATE VERDICT QUICK DCGM BURN COPIES DISK PAIRS ALL-REDUCE
gpu-01 Ready pass pass pass pass pass pass pass pass
gpu-02 Ready pass pass pass pass pass pass pass pass
gpu-03 Ready pass pass pass pass pass pass pass pass
gpu-04 Ready pass pass pass pass pass pass pass pass
gpu-05 Ready fail pass pass fail pass pass pass pass
gpu-06 Ready pass pass pass pass pass pass pass pass
gpu-07 Ready pass pass pass pass pass pass pass pass
gpu-08 Ready pass pass pass pass pass pass pass pass
What to fix (1)
! gpu-05 GPU 2: 76% of its peers' throughput under load (a straggler: every run it is in goes at its pace)
fix: Check its cooling and power (clocks held down), its PCIe link, and its NVLink; reset it
Pairs: 28 tested, best 392 Gb/s; bad: gpu-03–gpu-06
! gpu-03–gpu-06: 362 Gb/s (92% of the best): 362 Gb/s, 92% of the best pair's 392
All-reduce
MACHINES BUSBW GB/S EXPECTED SHARE VERDICT WHERE
2 361.8 368.0 98% pass one leaf
2 362.3 368.0 98% pass one leaf
4 357.4 368.0 97% pass one leaf
4 356.9 368.0 97% pass one leaf
8 291.4 312.8 93% pass across the spines
--machine <name> shows one machine's checks with every number
against its expected value; --json prints the report as the API
answers it; --html <file> saves the printable page. When the report
offers acceptance, it prints the astra astraeus validate command that
runs it.
$ curl -sS "$ASTRALYX_API/validation/report" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>" \
| jq '{counts, pairs: (.pairs | {best_gbps, lowest_latency_us, bad})}'
{
"counts": {
"machines": 8,
"passed": 7,
"warned": 0,
"failed": 1,
"inconclusive": 0,
"untested": 0,
"accepting": 0,
"observe_only": 0
},
"pairs": {
"best_gbps": 392.1,
"lowest_latency_us": 1.62,
"bad": [
"gpu-03–gpu-06"
]
}
}
The answer also has summary, suggestion (the offer, when there is
one), machines, pairs.cells, allreduce (the curve's points), the
10 most recent validations and the policies. Every field is in the
API reference.
In the example, gpu-05's GPU 2 is a straggler. The pair gpu-03–gpu-06
is slower than every other pair, also when tested alone, while each of its
rails met its own expected value: the fault is on the path between them —
a cable, a transceiver or a switch port — so it fails the validation and
is marked in the pairs' matrix, but neither machine fails by it. Test each
machine against other peers to tell which end, if either, is at fault.
Your AI assistant can read the report too, with the tool
get_validation_report
(Connect your AI assistant).
One machine's checks#
Open Compute → Machines and select the machine. Its badges show its quick check and its last acceptance; Checks shows each check's latest result, its numbers against what was expected, and what to fix.
Its line of the report, what to fix on it, then each check: its verdict, its summary, and every number against its expected value.
$ curl -sS "$ASTRALYX_API/machines/gpu-05/validation" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>" \
| jq '{name, observe_only, class, quick: .record.quick.summary, acceptance: .record.acceptance.verdict, verdict: .report.verdict}'
{
"name": "gpu-05",
"observe_only": false,
"class": "datacenter",
"quick": "gpu-05: 8 GPU(s) answer, driver 570.124.06, DCGM r1 passed",
"acceptance": "fail",
"verdict": "fail"
}
record holds the machine's badges — quick, acceptance, the latest
of each check (checks) — and gate while it waits to pass
acceptance; report is its line of the Cluster Report. The same
validation record comes with each machine of GET /machines.
Print or share it#
The printable page holds the summary, the counts, the offer, the machines by check, what to fix, every number against its expected value, the pairs' matrix and the all-reduce curve, with how expected values are derived. It is one HTML file, with no scripts, laid out for A4: print it to PDF from a browser.
Require acceptance in a pool#
A validation policy says, for one pool's machines, which checks may run
there and when, and whether new machines must pass acceptance before they
work. The pool * is the default for pools without a policy of their own
(and for machines without a pool); without any policy, the defaults below
apply. A policy belongs to the organisation that sets it.
With require_acceptance on:
- A machine of the pool that joins after the policy was turned on waits as Accepting: it takes no work but its checks. Machines that joined before are not held.
- Acceptance starts by itself, for the machines waiting together, 2 minutes after the last of them joined.
- A machine that passes (
passorwarn) takes work at once: eventAcceptance passed: <machine> takes work. - A machine that fails, or whose acceptance is inconclusive or cancelled, is Acceptance failed and still takes no work, until an acceptance run on it passes or an admin accepts it anyway.
Turning require_acceptance off releases the machines it holds. A run
that waits says why for each such machine: machine is being accepted: it
takes work once its checks pass (pool h100's policy: new machines must pass
acceptance), or machine failed acceptance and takes no work until it
passes (…).
| Field | Type | Default | Description |
|---|---|---|---|
pool |
string | the path's | The pool (pool=<name>): letters, digits, -, _, .; or *. Optional in the body: the path names it. |
require_acceptance |
boolean | false |
New machines must pass acceptance before they take work. |
require_acceptance_since |
time | — | Read-only: when require_acceptance was turned on. Kept while it stays on. |
checks |
list of check names | every check of the suite | The checks acceptance runs here: any of dcgm-diag, gpu-burn, nvbandwidth, fio, ib-write-bw, nccl-allreduce. It does not limit the network suite. |
burn_minutes |
integer, 1 to 240 | 10 |
The burn-in's length (deep: at least 30). |
dcgm_level |
integer, 1 to 4 | 2 |
DCGM's level for acceptance (deep: at least 3). |
max_concurrent_machines |
integer, at least 1 | 8 |
Most of the pool's machines under test at once. |
fabric_tests |
boolean | true |
Pairs and all-reduce at all. Off, acceptance checks each machine alone and the network suite has nothing to run. |
windows |
list of windows | none | When tests across the spines may run while other work runs. |
include_scale_to_zero |
boolean | false |
Also check machines that host a deployment scaled to zero. |
periodic |
{suite, every_hours} |
none | Kept with the policy; checks on a schedule come in a later release. |
updated_by, updated_at |
Read-only: who changed it last, and when. |
A validation is planned with the policy of its first machine (by name): check pools whose policies differ one pool at a time.
- Open Compute → Validation, then Policies.
- Select the pool's policy, or Add a policy and choose the pool
(
*for the default). - Turn on New machines must pass acceptance, and set the burn-in, DCGM's level, the machines at once and the windows as you need.
- Select Save.
astra does not change policies. Use the console or the API;
astra astraeus report --json shows them (policies).
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validation-policies/h100" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"require_acceptance": true, "max_concurrent_machines": 4,
"windows": [{"days": ["sat", "sun"], "start": "22:00", "end": "06:00", "timezone": "Europe/Lisbon"}]}' \
| jq -c '{pool, require_acceptance, require_acceptance_since, burn_minutes, max_concurrent_machines, windows}'
{"pool":"h100","require_acceptance":true,"require_acceptance_since":"2026-10-05T10:20:44.902117Z","burn_minutes":10,"max_concurrent_machines":4,"windows":[{"days":["sat","sun"],"start":"22:00","end":"06:00","timezone":"Europe/Lisbon"}]}
PUT replaces the pool's policy whole: a field left out takes its
default. DELETE …/validation-policies/h100 removes it (204; 404
POLICY_NOT_FOUND when there is none), and the default applies.
GET $ASTRALYX_API/validation-policies lists the policies, with the
defaults as default.
Changes are recorded in the audit log (cluster.validation.policy,
cluster.validation.policy_delete).
Windows#
A window lets tests cross the spines while other work runs: on its days
(mon to sun; none: every day), from start to end (HH:MM), in
timezone (an IANA name; default UTC). An end before its start runs
past midnight, into the next day: sat, sun 22:00–06:00 covers Saturday
22:00 to Sunday 06:00 and Sunday 22:00 to Monday 06:00.
Accept a machine anyway#
An organisation admin can let a machine that is Accepting or Acceptance failed take work without passing — for example a machine the vendor burned in, or one whose only failure you accept.
Open the machine and select Accept anyway. Say why; the reason is kept with the machine.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/nodes/gpu-09/accept" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"reason": "burned in by the vendor"}'
{"name":"gpu-09","accepted":true}
The body is optional. A machine that is not waiting answers
409 NOT_ACCEPTING.
The machine takes work at once. Its acceptance badge becomes warn
(accepted by <who> without passing: <reason>), and the machine has the
event Accepted by <who> after failing acceptance: it takes work
(<reason>) — or without waiting for acceptance, when it was still
Accepting. It is recorded in the audit log (cluster.node.accept).
Reference#
The request#
POST /organizations/{org}/clusters/{cluster}/validations takes:
| Field | Type | Default | Description |
|---|---|---|---|
suite |
quick, acceptance, network or deep |
quick |
The checks. |
machines |
list of machine names | every machine of the organisation on the cluster (of pool, when given) that is Up, and not in observe-only mode unless explicit |
The machines. Each must be Up (400 NODE_NOT_UP); one in observe-only mode needs explicit (400 OBSERVE_ONLY). |
pool |
string | none | Without machines: the pool's machines. |
explicit |
boolean | false |
An admin's explicit start: also machines in observe-only mode, cordoned machines, and machines where something outside Astraeus was seen using a GPU. |
options |
object | {} |
See below. |
note |
string | empty | Why; kept with the validation. |
| Option | Type | Default | Description |
|---|---|---|---|
burn_minutes |
integer | the policy's | The burn-in's length, kept within 1 to 240 (deep: at least 30). |
dcgm_level |
integer | the policy's | DCGM's level, kept within 1 to 4 (deep: at least 3). |
cross_spine |
boolean | false |
Test across the spines now, even with other work on the cluster. |
include_scale_to_zero |
boolean | the policy's | Also check machines that host a deployment scaled to zero. |
peers |
integer | 3 |
How many peers of the cluster each machine is paired with, when not every pair is tested. |
every_pair |
boolean | false |
Every pair of the machines, whatever the cluster's history and size. |
fabric_burn |
boolean | false |
A burn of the links: the pairs and all-reduce run ten times longer, loading the fabric for minutes. Only when asked; a policy never does. |
An unknown field, or a suite that does not exist, answers
400 INVALID_BODY.
What a validation and its steps say#
| Field | What |
|---|---|
id |
val-<date>-<6 hex digits>. |
reason |
Why it was made: on-demand (a person, through the console, the API or an assistant), acceptance-policy (a pool's policy), installer (--accept). |
requested_by |
Who asked: the person (console:<e-mail>), policy or installer. |
gates |
Its machines wait for it before they take work (a policy's acceptance). |
summary |
Once over, what it found, in a sentence: 7/8 machines pass; …. |
steps[].id |
The check and its machines: dcgm-diag.gpu-01, ib-write-bw.gpu-01.gpu-08; a retest ends in .retest1, .retest2… |
steps[].phase |
When it may run: 0 the machines' checks, 1 and up the pairs' rounds, 500 and up the retests, 1000 and up all-reduce inside a leaf group, 2000 and up across the spines. |
steps[].run |
The check run's name (chk-…) while it runs. |
steps[].params |
What the step is beyond its check, such as level, minutes, directory, round, machines, across: spine, halving, concurrency, burn (a burn of the links). |
steps[].waiting_since |
Since when it waits for its machines, its phase open: its 6-hour limit counts from there. |
steps[].result |
verdict, summary, measurements (each value against expected, in unit, with its verdict), problems (subject, what, fix), noise (other work on the links it crossed) and output (the end of the tool's output, up to 6 KiB). |
Events of a machine#
| Event | When |
|---|---|
Quick check: <what is wrong> |
A quick check warned or failed. |
Check <check> stopped for real work (preempted): it runs again when the machine is free |
A check was preempted. |
Acceptance passed: <machine> takes work |
A machine waiting for acceptance passed it. |
Acceptance failed: <what failed> |
A machine waiting for acceptance failed it, or it was inconclusive. |
Acceptance passed (<validation>), Acceptance failed: …, Acceptance inconclusive: … |
Acceptance of a machine that was not waiting for it. |
Accepted by <who> after failing acceptance: it takes work |
An admin accepted it (or without waiting for acceptance). |
Converted to normal by <who>: it takes work, Observe-only mode: set by <who> (it takes no work) |
See Observe-only mode. |
<who> is the person who acted, as console:<e-mail>. They are the
machine's events: on its page, among the cluster's events, and in event
streams' machines category
(Events, audit and event streams).
The API#
| Method and path | What | Who |
|---|---|---|
GET /validations |
Validations, newest first, without steps | viewer or above |
GET /validations/{id} |
One validation, with its steps | viewer or above |
GET /validation/report |
The Cluster Report (?format=html: the printable page) |
viewer or above |
GET /validation-policies |
The policies, and the defaults | viewer or above |
GET /machines/{name}/validation |
One machine's badges, latest results and line of the report | viewer or above |
POST /organizations/{org}/clusters/{cluster}/validations |
Start a suite | Organisation owner or admin |
POST /organizations/{org}/clusters/{cluster}/validations/{id}/cancel |
Cancel one | Organisation owner or admin |
PUT, DELETE /organizations/{org}/clusters/{cluster}/validation-policies/{pool} |
Set or remove a pool's policy | Organisation owner or admin |
POST /organizations/{org}/clusters/{cluster}/nodes/{machine}/accept |
Accept a machine without its passing | Organisation owner or admin |
PUT /organizations/{org}/clusters/{cluster}/nodes/{machine}/observe-only |
Convert a machine to normal, or make it observe-only (Observe-only mode) | Organisation owner or admin |
The paths in the first part act in a workspace and see the machines of its
pools: a validation of other machines answers 404 VALIDATION_NOT_FOUND.
Every field is in the Astraeus API reference
and the administration API reference.
Limits#
| Limit | Value |
|---|---|
| A check's time to stop when preempted | 5 s |
| Preemptions before a step is skipped | 3 |
| A check run not placed | Withdrawn after 90 s; its step waits again |
| A step's wait for its machines | 6 hours from when its phase opened and they were not free (again after each preemption); then skipped |
| Machines under test at once, per pool | The policy's max_concurrent_machines (8) |
Acceptance after machines join (a policy, --accept) |
2 minutes after the last of them joined |
| A quick check | 150 s for the answer, of which DCGM's part 90 s; tried again after 2 minutes when it could not be made |
| Every pair tested | Up to 64 new machines; more are paired with peers |
| Fabric nearly idle | Other work holds at most 5% of its GPUs, and no run spans machines |
| A tool's output kept with a result | Its last 6 KiB |
| Validations kept once over | The 100 most recent |
Troubleshooting#
A waiting step's why names the machine and the reason.
| Symptom | Cause | Fix |
|---|---|---|
A step waits: gpu-03: its GPUs are in use |
A run holds GPUs there. | Nothing: the check runs once the machine is free. Name other machines, or reserve these for maintenance (Reservations). |
gpu-03: serves inference replicas |
An Eos deployment has a replica there. Such machines are never checked. | Move or scale the deployment, then check the machine. |
gpu-03: GPU 2 is used by something outside Astraeus |
A program Astraeus did not start holds a GPU (GPUs). | Stop it, or start with explicit: true (the check still waits for that GPU to be free). |
gpu-03: hosts a deployment scaled to zero (…) |
A check could delay the deployment's wake-up. | Set include_scale_to_zero in the policy, or in the request's options. |
gpu-03: cordoned |
Cordoned machines are checked only on an explicit start. Their steps wait, and hold the next phase, until skipped. | Uncordon it, start with explicit: true, or name the machines to leave it out. |
gpu-03: in observe-only mode: only an admin's explicit run checks it |
See Observe-only mode. | Start with explicit: true. |
gpu-03: another check runs there |
One check at a time on a machine. | Nothing: it runs next. |
4 machine(s) of pool h100 under test already (the policy's most: 4) |
The pool's limit of machines at once. | Wait, or raise max_concurrent_machines. |
crosses the spines: waits for the policy's window, or the fabric nearly idle |
A pair under different leaf groups, or all-reduce across them, while other work runs. | Add a window, wait for a quiet fabric, or start with cross_spine: true. |
its run could not be placed: the machines got busy |
The machines took work between the check's look and its start. | Nothing: it waits for them again. |
A step Skipped: preempted 3 times: real work kept needing its machines |
The machines are in steady demand. | Run checks in a quiet hour, or reserve the machines for maintenance. |
A step Preempted: stopped for real work (2 times); runs again when its machines are free — gpu-03: its GPUs are in use |
Work took the machines while the check ran; it waits for them again. | Nothing: it runs again when they are free, unless preempted 3 times. |
A step Skipped: its machines were not free within 6 h: … |
Once its phase opened, its machines stayed busy for 6 hours. | As above; check fewer machines at once. |
A result inconclusive: … not judged: other work loaded the links it crossed (…) |
Another run crossed the same spines during the test. | Run it again in a window, or with the fabric idle. |
DCGM's diagnostic could not run: … (inconclusive) |
The DCGM image could not be pulled, or the driver did not answer it. | Let the machine pull nvcr.io/nvidia/cloud-native/dcgm; check the driver (nvidia-smi). |
The burn-in could not run: …, nvbandwidth could not run: …, fio could not run: …, ib_write_bw could not run: … |
The checks image could not be pulled, or the tool failed; the summary gives its exit and reason. | Let the machine pull ghcr.io/astralyx-cloud/astraeus-checks; read the step's output. |
A GPU check fails at once: gpu-03 GPU 5: fenced: … |
The GPU is fenced for a fault: a check cannot have it. | Follow the fault's action, then clear the fault. |
| A machine has no Data disk result | It has no data location. | Choose one on the machine's page (Drives and data). |
The validation ends Inconclusive: Nothing to check: … |
No check applies: the network suite on machines without RDMA ports or GPUs, or fabric tests turned off. | Check the machines and the policy. |
400 OBSERVE_ONLY |
A machine named is in observe-only mode. | Add "explicit": true. |
400 NODE_NOT_UP |
A machine named is down or not reporting. | Wait for it to be Up. |
400 NO_MACHINES |
Every machine was left out: down, or in observe-only mode. | Name machines that are Up, or add explicit. |
400 INVALID_VALIDATION (a policy) |
A value out of range: dcgm_level 9: 1 to 4, burn_minutes 300: 1 to 240, a window's window time "25:00": HH:MM. |
Correct the value. |
403 ORG_ADMIN_REQUIRED |
Only organisation owners and admins start, cancel and accept, convert machines, and change policies — in the console, with astra, or through the API. |
Ask one. |
A new machine takes no work: machine is being accepted: … |
Its pool requires acceptance, which has not passed yet. | Wait for it; follow it in Compute → Validation. |
machine failed acceptance and takes no work until it passes (…) |
Its acceptance failed, was inconclusive or was cancelled. | Fix what it found and run acceptance again, or accept it anyway. |