Adopt Astraeus next to Slurm or Kubernetes#
You run a Slurm partition or a Kubernetes node pool today, and you want to move to Astraeus without a cutover weekend. Install the agent in observe-only mode beside the scheduler that runs the machines now: it reads everything — GPU health, topology, the Cluster Report — and places no work. When you are ready, accept the machines in a maintenance window and convert them, one partition at a time, while the rest keeps running where it is.
Before you begin#
- You are an owner or admin of the organisation: only organisation admins create join tokens and convert machines.
- The machines meet the requirements and already run Slurm or Kubernetes workloads you do not want disturbed yet.
- An organisation admin's API token,
curlandjq. Below,<org>and<cluster>stand for your organisation's and cluster's short names, andASTRALYX_APIishttps://api.astralyx.cloud/v1. - The example: a Slurm partition
slurm-aof four machines,gpu-21togpu-24.
1. Install in observe-only mode#
Add --observe-only and --no-data-dir to the installer's line, with a
pool of its own (one per partition or node pool), on each machine of
slurm-a:
$ echo '9c4e…' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --control-plane https://connect.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme Research' --org-id 0192f0c4-7d1e-7a51-9c33-5e8b2a4f6d10 --name gpu-21 --no-data-dir --observe-only && rm -f ./astraeus-token
Mint each machine's token with {"labels": {"pool": "slurm-a"}}, as in
Add a machine, or add --observe-only
--no-data-dir to the shell task of the Ansible playbook in Bring up a new
GPU cluster
to do all four at once. Nothing changes for the jobs Slurm is running
there: no container runs on these machines until an admin starts a check
explicitly, or converts them.
2. Give a workspace the pool, to evaluate it#
Grant pool=slurm-a to the workspace of the people who will look at what
Astraeus sees, so they see exactly these machines and nothing else of the
cluster (see Quotas, pools and terms).
Organisation admins see every machine of the cluster regardless of pool.
3. Read what Astraeus sees#
Compute → Machines lists gpu-21 … gpu-24, Up, with the
badge Observe-only. Each shows its GPU health, faults, and the
other scheduler's use of its GPUs as Used outside Astraeus.
Compute → Topology shows fabrics, leaf and spine groups and
findings; Compute → Validation shows them in the Cluster Report
with the state Observe-only.
4. Fix what topology finds, and export topology.conf#
Topology discovery needs nothing to start, and finds miswired cables,
degraded links and narrow PCIe links whether or not Astraeus runs anything.
Fix what it finds now — reseating a cable costs the same whichever
scheduler eventually benefits — and, while Slurm still runs the machines,
export the same topology as topology.conf for it:
Save it as Slurm's topology.conf, set TopologyPlugin=topology/tree (or
topology/block) in slurm.conf, and restart slurmctld. Your scheduler
now places jobs with the same rail and leaf-group awareness Astraeus uses,
before you have moved a single job. See
Export to Slurm.
5. Accept, in a maintenance window#
Checks never take a GPU another program holds, but the burn-in loads the machine for minutes and the fabric checks load its network. Drain the partition first:
(On Kubernetes: kubectl cordon <node> then kubectl drain <node>
--ignore-daemonsets --delete-emptydir-data.)
Compute → Validation → Run checks, suite Acceptance, machines
gpu-21 … gpu-24, turn on Check observe-only machines too, then
Start.
Fix what fails, and run it again until the four pass. explicit: true
(--explicit with astra, Check observe-only machines too in the
console) is required: without it, a request naming an observe-only
machine is refused (400 OBSERVE_ONLY).
6. Convert the partition#
Open each machine and select Convert to normal.
Each machine takes work from then on, like any machine of its pool; its quick check runs within moments, and it waits as Accepting only if the pool's policy requires acceptance and it has not passed yet (it has).
7. Move the work#
Submit new work as runs, or keep existing batch
scripts unchanged with astra slurm sbatch — it reads #SBATCH
directives, including --partition, which maps to the pool:
See Slurm compatibility for srun, squeue,
scancel, sinfo, and which #SBATCH options map to what.
8. Keep the machines out of the other scheduler#
Astraeus leaves alone a GPU another program already holds, but Slurm and Kubernetes do not know of Astraeus's work: two schedulers must never place work on the same GPUs. Leave the machines drained in Slurm (or cordoned in Kubernetes), or take them out of its configuration entirely, once they are converted.
9. Repeat for the next partition#
Go back to step 5 for the next partition or node pool — slurm-b,
gpu-pool-2, and so on — until nothing is left running under the old
scheduler.
What you get#
- A live read of your current fleet's health and topology, weeks before anything moves.
topology.conffor the scheduler you still run, from day one.- A partition-by-partition cutover with no window where the machines take no work at all.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
400 OBSERVE_ONLY starting checks |
Checks on an observe-only machine need an explicit start. | Add "explicit": true (--explicit, Check observe-only machines too). |
An explicit check keeps waiting: its run could not be placed: the machines got busy |
Slurm or Kubernetes jobs still hold GPUs there. | Drain the machine in Slurm or Kubernetes first. |
| Runs never land on a converted machine | Rerunning the installer did not change its mode: an admin's choice in Astraeus wins over the installer's flag. | astra astraeus observe-only <machine> off, or convert it in the console. |
astra: not used on Astraeus: --mail-type=END |
sbatch directives Astraeus does not use are named and dropped. |
Nothing to fix; the run still runs. |
A script calls srun inside the batch script |
There is no srun inside the container; each worker runs the whole script once. |
Drop srun from the script's commands. |
| GPUs still show Used outside Astraeus after converting | The machine was never removed from Slurm's or Kubernetes' configuration. | Remove it from the other scheduler's node list. |