Bring up a new GPU cluster#
A new rack of GPU servers arrives. You install the agent on all of them in one pass, put them in a pool, let Astraeus check every GPU, the data disk and the fabric between them, and fix whatever the Cluster Report finds — before a training run finds it for you. This recipe does that for eight 8-GPU machines on one InfiniBand fabric; the same steps work for any number.
Before you begin#
- You are an owner or admin of the organisation that owns the cluster: only organisation admins create join tokens and start acceptance.
- The machines meet the requirements: Linux with systemd on x86-64 or ARM64, root access, and outbound HTTPS to the cluster; an inventory and SSH access for Ansible, or a way to run one command as root on each.
- An organisation admin's API token,
curl,jqandansible-playbook. Below,<org>and<cluster>stand for your organisation's and cluster's short names, andASTRALYX_APIishttps://api.astralyx.cloud/v1. - The example: eight machines,
gpu-01togpu-08, each with eight H100 GPUs and eight 400 Gb/s InfiniBand ports, all on one fabric.
1. Install the agent on every machine at once#
Mint one join token per machine (a token is single-use) and run the
installer on all of them with Ansible, labelling them into the pool h100
as they join. This is the playbook from
Add a machine, unchanged:
- name: Add machines to Astraeus
hosts: gpu_servers
become: true
vars:
astraeus_console: https://console.astralyx.cloud
astraeus_org: <org>
astraeus_cluster: <cluster>
astraeus_agent_url: https://connect.astralyx.cloud
astraeus_org_name: <org name>
astraeus_org_id: <org id>
astraeus_releases: https://console.astralyx.cloud/releases
astraeus_pool: h100
astraeus_data_dir: /mnt/nvme0/astraeus
tasks:
- name: Is the machine already connected?
ansible.builtin.stat:
path: /etc/astraeus/token
register: astraeus_installed
- name: Create a join token for this host
ansible.builtin.uri:
url: "{{ astraeus_console }}/api/v1/orgs/{{ astraeus_org }}/clusters/{{ astraeus_cluster }}/enrollment-tokens"
method: POST
headers:
Authorization: "Bearer {{ lookup('ansible.builtin.env', 'ASTRAEUS_TOKEN') }}"
body_format: json
body: { labels: { pool: "{{ astraeus_pool }}" }, ttl_seconds: 3600 }
status_code: 201
delegate_to: localhost
become: false
register: enrollment
no_log: true
when: not astraeus_installed.stat.exists
- name: Copy the token to the host
ansible.builtin.copy:
content: "{{ enrollment.json.token }}"
dest: /root/astraeus-token
owner: root
mode: "0600"
no_log: true
when: not astraeus_installed.stat.exists
- name: Install the Astraeus agent
ansible.builtin.shell: >-
set -o pipefail &&
curl -fsSL {{ astraeus_console }}/api/v1/install.sh | sh -s --
--control-plane {{ astraeus_agent_url }}
--token-file /root/astraeus-token
--releases {{ astraeus_releases }}
--org '{{ astraeus_org_name }}' --org-id {{ astraeus_org_id }}
--name {{ inventory_hostname_short }}
--data-dir {{ astraeus_data_dir }}
--no-move
args: { executable: /bin/bash }
when: not astraeus_installed.stat.exists
- name: Remove the token file
ansible.builtin.file: { path: /root/astraeus-token, state: absent }
$ export ASTRALYX_TOKEN=$(cat ~/.astraeus-api-token)
$ ansible-playbook -i inventory.ini astraeus-machines.yml
A server with InfiniBand or RoCE ports keeps the host network by itself
(what NCCL over RDMA needs); the installer says so as it runs. Machines
with an NVIDIA GPU but no driver get one where the distribution has an
official way; add --install-nvidia-driver to the shell task's command
line to allow adding a repository without a terminal to ask.
2. Check that they all joined#
Open Compute → Machines: all eight appear Up within a few seconds of the installer finishing, with their GPUs, CPU, memory and pool filled in from their first reports.
Each machine checks itself within moments of coming Up (the quick check): a bad driver, a GPU off the bus or an unanswering NVML shows at once as a badge, long before acceptance.
3. Run acceptance#
Acceptance runs the GPU diagnostics, a 10-minute burn-in, GPU copy bandwidth and the data disk on each machine, then the fabric between every pair and an all-reduce across growing groups — about 32 minutes for eight machines. It never competes with real work: it starts only while the GPUs are free and gives way within seconds to anything that needs them (see How checks stay out of your work's way).
Open Compute → Validation. With no machine accepted yet, the report
offers it: select Run acceptance in the offer, or Run checks,
choose the suite Acceptance and the pool h100, then Start.
$ astra astraeus validate --suite acceptance --pool h100 --wait
validation val-20261006-7f3a2c: acceptance on gpu-01, gpu-02, gpu-03, gpu-04, gpu-05, gpu-06, gpu-07, gpu-08
https://console.astralyx.cloud/o/acme/w/vision/validation?cluster=lab-a&validation=val-20261006-7f3a2c
…
val-20261006-7f3a2c acceptance Passed (on-demand)
8/8 machines pass.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"suite": "acceptance", "pool": "h100", "note": "new rack, first bring-up"}' \
| jq -r .id
val-20261006-7f3a2c
Poll GET $ASTRALYX_API/validations/val-20261006-7f3a2c for its
state and steps, or follow it in the console.
Instead of running it by hand every time, add --accept to the install
command (or the Ansible shell task): acceptance starts by itself once a
machine's GPUs are free, 2 minutes after the last of a joining batch, and
never keeps the machine from work meanwhile. See
At install: --accept.
4. Read the Cluster Report#
The report puts every check's latest result together: a one-sentence summary, counts, each machine against each check, what to fix, the pairs' matrix and the all-reduce curve.
Compute → Validation is the report: the summary, the machines by check, what to fix, the pairs' matrix, the all-reduce curve. Download saves it as a printable page or JSON.
$ astra astraeus report
7/8 machines pass; gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace; pair gpu-03–gpu-06 at 92% of the best pair; all-reduce over 8 machines at 93% of expected (291 of 313 GB/s).
MACHINE STATE VERDICT QUICK DCGM BURN COPIES DISK PAIRS ALL-REDUCE
gpu-01 Ready pass pass pass pass pass pass pass pass
…
gpu-06 Ready warn pass pass pass pass pass pass pass
What to fix (1)
! gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace
fix: Reseat or replace the cable and transceiver of gpu-06 mlx5_5 (`mlxlink -d mlx5_5 -m` shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.
--html cluster-report.html saves the printable page.
5. Act on topology's findings#
Discovery works out the cluster's fabrics, leaf and spine groups from
what each machine reports, and flags what looks miswired or degraded — a
cable on the wrong rail's leaf, a port trained below its rate, a PCIe link
narrower than it should be. The report above quotes one: a degraded cable
on gpu-06's rail 5.
Open Compute → Topology → Findings for every finding, its evidence and its fix; open the machine for the same, beside its ports and their cabling.
$ astra astraeus topology
Findings (1)
! gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace
fix: Reseat or replace the cable and transceiver of gpu-06 mlx5_5 (`mlxlink -d mlx5_5 -m` shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.
After reseating the cable and transceiver, run the network suite alone to confirm it, without redoing the whole burn-in:
Every finding is explained in Topology, with what it means and what sets it right.
6. Require acceptance for machines that join later#
So the next batch of machines is checked the same way without you remembering to ask, set the pool's policy:
Compute → Validation → Policies, the pool h100, turn on New
machines must pass acceptance, set the burn-in and DCGM's level, and
Save.
A machine that joins the pool from now on waits as Accepting and takes no work but its checks, until it passes. See Require acceptance in a pool.
What you get#
- Eight machines, Up, in the pool
h100, every GPU, the data disk and the InfiniBand fabric checked against what the hardware should give. - A Cluster Report you can print or hand to whoever racked the servers, with every problem and its fix.
- A policy that holds the next batch of machines to the same bar before they take work.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
| A machine never appears Up | The installer could not reach the control plane, or the token expired (24 h from the console, 1 h by default from the API). | Check outbound HTTPS from the machine; mint a new token. |
ENROLLMENT_USED |
The token was already claimed by another machine. | Mint one token per machine; Ansible's playbook does this per host. |
A step waits: gpu-03: another check runs there |
One check at a time on a machine. | Nothing: it runs next. |
| Acceptance fails on one machine only | A real hardware problem: read What to fix on the report. | Follow the fix; accept it anyway with a reason if you choose to ship it as-is (astra astraeus accept gpu-03 --reason "…"). |
gpu-03: in observe-only mode |
The machine was installed with --observe-only. |
See Adopt Astraeus next to Slurm or Kubernetes if that was intentional; otherwise convert it. |
| All-reduce passes but is lower than expected | Check Topology for a finding on the machines involved before assuming the fabric is fine. |