Skip to content

Adopt Astraeus next to Slurm or Kubernetes#

You run a Slurm partition or a Kubernetes node pool today, and you want to move to Astraeus without a cutover weekend. Install the agent in observe-only mode beside the scheduler that runs the machines now: it reads everything — GPU health, topology, the Cluster Report — and places no work. When you are ready, accept the machines in a maintenance window and convert them, one partition at a time, while the rest keeps running where it is.

Before you begin#

  • You are an owner or admin of the organisation: only organisation admins create join tokens and convert machines.
  • The machines meet the requirements and already run Slurm or Kubernetes workloads you do not want disturbed yet.
  • An organisation admin's API token, curl and jq. Below, <org> and <cluster> stand for your organisation's and cluster's short names, and ASTRALYX_API is https://api.astralyx.cloud/v1.
  • The example: a Slurm partition slurm-a of four machines, gpu-21 to gpu-24.

1. Install in observe-only mode#

Add --observe-only and --no-data-dir to the installer's line, with a pool of its own (one per partition or node pool), on each machine of slurm-a:

$ echo '9c4e…' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --control-plane https://connect.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme Research' --org-id 0192f0c4-7d1e-7a51-9c33-5e8b2a4f6d10 --name gpu-21 --no-data-dir --observe-only && rm -f ./astraeus-token

Mint each machine's token with {"labels": {"pool": "slurm-a"}}, as in Add a machine, or add --observe-only --no-data-dir to the shell task of the Ansible playbook in Bring up a new GPU cluster to do all four at once. Nothing changes for the jobs Slurm is running there: no container runs on these machines until an admin starts a check explicitly, or converts them.

2. Give a workspace the pool, to evaluate it#

Grant pool=slurm-a to the workspace of the people who will look at what Astraeus sees, so they see exactly these machines and nothing else of the cluster (see Quotas, pools and terms). Organisation admins see every machine of the cluster regardless of pool.

3. Read what Astraeus sees#

Compute → Machines lists gpu-21 … gpu-24, Up, with the badge Observe-only. Each shows its GPU health, faults, and the other scheduler's use of its GPUs as Used outside Astraeus. Compute → Topology shows fabrics, leaf and spine groups and findings; Compute → Validation shows them in the Cluster Report with the state Observe-only.

$ astra astraeus machines
MACHINE  STATE  GPUS                       CPU   POOL
gpu-21   Up     8× NVIDIA H100 80GB HBM3    224   slurm-a
$ astra astraeus report --machine gpu-21
$ curl -sS "$ASTRALYX_API/machines/gpu-21/validation" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
    -H "Astralyx-Workspace: <org>/<workspace>" \
  | jq '{name, observe_only, state: .report.state}'
{"name": "gpu-21", "observe_only": true, "state": "Observe-only"}

4. Fix what topology finds, and export topology.conf#

Topology discovery needs nothing to start, and finds miswired cables, degraded links and narrow PCIe links whether or not Astraeus runs anything. Fix what it finds now — reseating a cable costs the same whichever scheduler eventually benefits — and, while Slurm still runs the machines, export the same topology as topology.conf for it:

Compute → Topology → Export for Slurm, choose Tree or Block, and Download topology.conf.

$ astra astraeus topology --slurm tree > topology.conf

Save it as Slurm's topology.conf, set TopologyPlugin=topology/tree (or topology/block) in slurm.conf, and restart slurmctld. Your scheduler now places jobs with the same rail and leaf-group awareness Astraeus uses, before you have moved a single job. See Export to Slurm.

5. Accept, in a maintenance window#

Checks never take a GPU another program holds, but the burn-in loads the machine for minutes and the fabric checks load its network. Drain the partition first:

$ scontrol update NodeName=gpu-[21-24] State=DRAIN Reason="Astraeus acceptance"

(On Kubernetes: kubectl cordon <node> then kubectl drain <node> --ignore-daemonsets --delete-emptydir-data.)

Compute → Validation → Run checks, suite Acceptance, machines gpu-21 … gpu-24, turn on Check observe-only machines too, then Start.

$ astra astraeus validate --suite acceptance --machines gpu-21,gpu-22,gpu-23,gpu-24 --explicit --wait
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"suite": "acceptance", "machines": ["gpu-21", "gpu-22", "gpu-23", "gpu-24"], "explicit": true, "note": "slurm-a, drained"}'

Fix what fails, and run it again until the four pass. explicit: true (--explicit with astra, Check observe-only machines too in the console) is required: without it, a request naming an observe-only machine is refused (400 OBSERVE_ONLY).

6. Convert the partition#

Open each machine and select Convert to normal.

$ for m in gpu-21 gpu-22 gpu-23 gpu-24; do astra astraeus observe-only $m off; done
$ for m in gpu-21 gpu-22 gpu-23 gpu-24; do
    curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/$m/observe-only" \
      -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' -d '{"observe_only": false}'
  done

Each machine takes work from then on, like any machine of its pool; its quick check runs within moments, and it waits as Accepting only if the pool's policy requires acceptance and it has not passed yet (it has).

7. Move the work#

Submit new work as runs, or keep existing batch scripts unchanged with astra slurm sbatch — it reads #SBATCH directives, including --partition, which maps to the pool:

$ astra slurm sbatch --partition=slurm-a train.sbatch
Submitted batch job train

See Slurm compatibility for srun, squeue, scancel, sinfo, and which #SBATCH options map to what.

8. Keep the machines out of the other scheduler#

Astraeus leaves alone a GPU another program already holds, but Slurm and Kubernetes do not know of Astraeus's work: two schedulers must never place work on the same GPUs. Leave the machines drained in Slurm (or cordoned in Kubernetes), or take them out of its configuration entirely, once they are converted.

9. Repeat for the next partition#

Go back to step 5 for the next partition or node pool — slurm-b, gpu-pool-2, and so on — until nothing is left running under the old scheduler.

What you get#

  • A live read of your current fleet's health and topology, weeks before anything moves.
  • topology.conf for the scheduler you still run, from day one.
  • A partition-by-partition cutover with no window where the machines take no work at all.

Troubleshooting#

Symptom Cause Fix
400 OBSERVE_ONLY starting checks Checks on an observe-only machine need an explicit start. Add "explicit": true (--explicit, Check observe-only machines too).
An explicit check keeps waiting: its run could not be placed: the machines got busy Slurm or Kubernetes jobs still hold GPUs there. Drain the machine in Slurm or Kubernetes first.
Runs never land on a converted machine Rerunning the installer did not change its mode: an admin's choice in Astraeus wins over the installer's flag. astra astraeus observe-only <machine> off, or convert it in the console.
astra: not used on Astraeus: --mail-type=END sbatch directives Astraeus does not use are named and dropped. Nothing to fix; the run still runs.
A script calls srun inside the batch script There is no srun inside the container; each worker runs the whole script once. Drop srun from the script's commands.
GPUs still show Used outside Astraeus after converting The machine was never removed from Slurm's or Kubernetes' configuration. Remove it from the other scheduler's node list.