Skip to content

Bring up a new GPU cluster#

A new rack of GPU servers arrives. You install the agent on all of them in one pass, put them in a pool, let Astraeus check every GPU, the data disk and the fabric between them, and fix whatever the Cluster Report finds — before a training run finds it for you. This recipe does that for eight 8-GPU machines on one InfiniBand fabric; the same steps work for any number.

Before you begin#

  • You are an owner or admin of the organisation that owns the cluster: only organisation admins create join tokens and start acceptance.
  • The machines meet the requirements: Linux with systemd on x86-64 or ARM64, root access, and outbound HTTPS to the cluster; an inventory and SSH access for Ansible, or a way to run one command as root on each.
  • An organisation admin's API token, curl, jq and ansible-playbook. Below, <org> and <cluster> stand for your organisation's and cluster's short names, and ASTRALYX_API is https://api.astralyx.cloud/v1.
  • The example: eight machines, gpu-01 to gpu-08, each with eight H100 GPUs and eight 400 Gb/s InfiniBand ports, all on one fabric.

1. Install the agent on every machine at once#

Mint one join token per machine (a token is single-use) and run the installer on all of them with Ansible, labelling them into the pool h100 as they join. This is the playbook from Add a machine, unchanged:

astraeus-machines.yml
- name: Add machines to Astraeus
  hosts: gpu_servers
  become: true
  vars:
    astraeus_console: https://console.astralyx.cloud
    astraeus_org: <org>
    astraeus_cluster: <cluster>
    astraeus_agent_url: https://connect.astralyx.cloud
    astraeus_org_name: <org name>
    astraeus_org_id: <org id>
    astraeus_releases: https://console.astralyx.cloud/releases
    astraeus_pool: h100
    astraeus_data_dir: /mnt/nvme0/astraeus
  tasks:
    - name: Is the machine already connected?
      ansible.builtin.stat:
        path: /etc/astraeus/token
      register: astraeus_installed

    - name: Create a join token for this host
      ansible.builtin.uri:
        url: "{{ astraeus_console }}/api/v1/orgs/{{ astraeus_org }}/clusters/{{ astraeus_cluster }}/enrollment-tokens"
        method: POST
        headers:
          Authorization: "Bearer {{ lookup('ansible.builtin.env', 'ASTRAEUS_TOKEN') }}"
        body_format: json
        body: { labels: { pool: "{{ astraeus_pool }}" }, ttl_seconds: 3600 }
        status_code: 201
      delegate_to: localhost
      become: false
      register: enrollment
      no_log: true
      when: not astraeus_installed.stat.exists

    - name: Copy the token to the host
      ansible.builtin.copy:
        content: "{{ enrollment.json.token }}"
        dest: /root/astraeus-token
        owner: root
        mode: "0600"
      no_log: true
      when: not astraeus_installed.stat.exists

    - name: Install the Astraeus agent
      ansible.builtin.shell: >-
        set -o pipefail &&
        curl -fsSL {{ astraeus_console }}/api/v1/install.sh | sh -s --
        --control-plane {{ astraeus_agent_url }}
        --token-file /root/astraeus-token
        --releases {{ astraeus_releases }}
        --org '{{ astraeus_org_name }}' --org-id {{ astraeus_org_id }}
        --name {{ inventory_hostname_short }}
        --data-dir {{ astraeus_data_dir }}
        --no-move
      args: { executable: /bin/bash }
      when: not astraeus_installed.stat.exists

    - name: Remove the token file
      ansible.builtin.file: { path: /root/astraeus-token, state: absent }
$ export ASTRALYX_TOKEN=$(cat ~/.astraeus-api-token)
$ ansible-playbook -i inventory.ini astraeus-machines.yml

A server with InfiniBand or RoCE ports keeps the host network by itself (what NCCL over RDMA needs); the installer says so as it runs. Machines with an NVIDIA GPU but no driver get one where the distribution has an official way; add --install-nvidia-driver to the shell task's command line to allow adding a repository without a terminal to ask.

2. Check that they all joined#

Open Compute → Machines: all eight appear Up within a few seconds of the installer finishing, with their GPUs, CPU, memory and pool filled in from their first reports.

$ astra astraeus machines
MACHINE   STATE  GPUS                        CPU   POOL
gpu-01    Up     8× NVIDIA H100 80GB HBM3     224   h100
gpu-02    Up     8× NVIDIA H100 80GB HBM3     224   h100
…
gpu-08    Up     8× NVIDIA H100 80GB HBM3     224   h100
$ curl -sS "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" \
  | jq -r '.items[] | "\(.metadata.name)\t\(.status.state)\t\(.metadata.labels.pool)"'
gpu-01  Up  h100
…

Each machine checks itself within moments of coming Up (the quick check): a bad driver, a GPU off the bus or an unanswering NVML shows at once as a badge, long before acceptance.

3. Run acceptance#

Acceptance runs the GPU diagnostics, a 10-minute burn-in, GPU copy bandwidth and the data disk on each machine, then the fabric between every pair and an all-reduce across growing groups — about 32 minutes for eight machines. It never competes with real work: it starts only while the GPUs are free and gives way within seconds to anything that needs them (see How checks stay out of your work's way).

Open Compute → Validation. With no machine accepted yet, the report offers it: select Run acceptance in the offer, or Run checks, choose the suite Acceptance and the pool h100, then Start.

$ astra astraeus validate --suite acceptance --pool h100 --wait
validation val-20261006-7f3a2c: acceptance on gpu-01, gpu-02, gpu-03, gpu-04, gpu-05, gpu-06, gpu-07, gpu-08
https://console.astralyx.cloud/o/acme/w/vision/validation?cluster=lab-a&validation=val-20261006-7f3a2c
…
val-20261006-7f3a2c  acceptance  Passed  (on-demand)
8/8 machines pass.
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"suite": "acceptance", "pool": "h100", "note": "new rack, first bring-up"}' \
  | jq -r .id
val-20261006-7f3a2c

Poll GET $ASTRALYX_API/validations/val-20261006-7f3a2c for its state and steps, or follow it in the console.

Instead of running it by hand every time, add --accept to the install command (or the Ansible shell task): acceptance starts by itself once a machine's GPUs are free, 2 minutes after the last of a joining batch, and never keeps the machine from work meanwhile. See At install: --accept.

4. Read the Cluster Report#

The report puts every check's latest result together: a one-sentence summary, counts, each machine against each check, what to fix, the pairs' matrix and the all-reduce curve.

Compute → Validation is the report: the summary, the machines by check, what to fix, the pairs' matrix, the all-reduce curve. Download saves it as a printable page or JSON.

$ astra astraeus report
7/8 machines pass; gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace; pair gpu-03–gpu-06 at 92% of the best pair; all-reduce over 8 machines at 93% of expected (291 of 313 GB/s).

MACHINE  STATE  VERDICT  QUICK  DCGM  BURN  COPIES  DISK  PAIRS  ALL-REDUCE
gpu-01   Ready  pass     pass   pass  pass  pass    pass  pass   pass
…
gpu-06   Ready  warn     pass   pass  pass  pass    pass  pass   pass

What to fix (1)
! gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace
    fix: Reseat or replace the cable and transceiver of gpu-06 mlx5_5 (`mlxlink -d mlx5_5 -m` shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.

--html cluster-report.html saves the printable page.

$ curl -sS "$ASTRALYX_API/validation/report" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
    -H "Astralyx-Workspace: <org>/<workspace>" | jq '{summary, counts}'

5. Act on topology's findings#

Discovery works out the cluster's fabrics, leaf and spine groups from what each machine reports, and flags what looks miswired or degraded — a cable on the wrong rail's leaf, a port trained below its rate, a PCIe link narrower than it should be. The report above quotes one: a degraded cable on gpu-06's rail 5.

Open Compute → Topology → Findings for every finding, its evidence and its fix; open the machine for the same, beside its ports and their cabling.

$ astra astraeus topology
Findings (1)
  ! gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace
      fix: Reseat or replace the cable and transceiver of gpu-06 mlx5_5 (`mlxlink -d mlx5_5 -m` shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.
$ curl -sS "$ASTRALYX_API/machines/gpu-06/topology" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
    -H "Astralyx-Workspace: <org>/<workspace>" | jq '.findings'

After reseating the cable and transceiver, run the network suite alone to confirm it, without redoing the whole burn-in:

$ astra astraeus validate --suite network --machines gpu-06 --wait

Every finding is explained in Topology, with what it means and what sets it right.

6. Require acceptance for machines that join later#

So the next batch of machines is checked the same way without you remembering to ask, set the pool's policy:

Compute → Validation → Policies, the pool h100, turn on New machines must pass acceptance, set the burn-in and DCGM's level, and Save.

$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validation-policies/h100" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"require_acceptance": true, "max_concurrent_machines": 8}'

A machine that joins the pool from now on waits as Accepting and takes no work but its checks, until it passes. See Require acceptance in a pool.

What you get#

  • Eight machines, Up, in the pool h100, every GPU, the data disk and the InfiniBand fabric checked against what the hardware should give.
  • A Cluster Report you can print or hand to whoever racked the servers, with every problem and its fix.
  • A policy that holds the next batch of machines to the same bar before they take work.

Troubleshooting#

Symptom Cause Fix
A machine never appears Up The installer could not reach the control plane, or the token expired (24 h from the console, 1 h by default from the API). Check outbound HTTPS from the machine; mint a new token.
ENROLLMENT_USED The token was already claimed by another machine. Mint one token per machine; Ansible's playbook does this per host.
A step waits: gpu-03: another check runs there One check at a time on a machine. Nothing: it runs next.
Acceptance fails on one machine only A real hardware problem: read What to fix on the report. Follow the fix; accept it anyway with a reason if you choose to ship it as-is (astra astraeus accept gpu-03 --reason "…").
gpu-03: in observe-only mode The machine was installed with --observe-only. See Adopt Astraeus next to Slurm or Kubernetes if that was intentional; otherwise convert it.
All-reduce passes but is lower than expected Check Topology for a finding on the machines involved before assuming the fabric is fine.