Skip to content

Slurm compatibility#

If your team has Slurm batch scripts, you can run them on Astraeus without rewriting them as run specifications. The astra CLI provides Slurm's commands — sbatch, srun, squeue, scancel, sinfo — and turns each script's #SBATCH directives into a run. Inside every worker, Slurm's environment variables are set, so scripts and libraries that read them keep working.

Slurm's words map to Astraeus' as: a job is a run, a task or node is a worker on a machine, a partition is a pool.

Before you begin#

  • astra is installed and signed in, and a workspace is chosen (astra use <org>/<workspace>). See Install the CLI.
  • Astraeus runs everything in containers. Name the image with Pyxis' #SBATCH --container-image=<image>, or set ASTRA_IMAGE. Your code must be in the image or on a drive the image can reach: sbatch sends only the script.

Use the commands#

Run them as subcommands of astra slurm:

$ astra slurm sbatch train.sbatch
Submitted batch job llama

Or install them under their own names, as links to astra, in a directory on your PATH:

$ astra slurm install-shims ~/.local/bin
/home/ada/.local/bin/sbatch → /usr/local/bin/astra
/home/ada/.local/bin/srun → /usr/local/bin/astra
/home/ada/.local/bin/squeue → /usr/local/bin/astra
/home/ada/.local/bin/scancel → /usr/local/bin/astra
/home/ada/.local/bin/sinfo → /usr/local/bin/astra
$ sbatch train.sbatch
Submitted batch job llama

Submit a batch script#

train.sbatch
#!/bin/bash
#SBATCH --job-name=llama
#SBATCH --nodes=2
#SBATCH --gres=gpu:8
#SBATCH --cpus-per-task=64
#SBATCH --mem=512G
#SBATCH --time=2-00:00:00
#SBATCH --partition=h100
#SBATCH --container-image=registry.example.com/nlp/train:2026-09

torchrun --nnodes="$SLURM_NNODES" --node-rank="$SLURM_PROCID" \
  --nproc-per-node="$SLURM_GPUS_ON_NODE" \
  --master-addr="$MASTER_ADDR" --master-port="$MASTER_PORT" \
  /opt/train/train.py
$ astra slurm sbatch train.sbatch
Submitted batch job llama

What happens:

  1. The #SBATCH lines are read, then the options given on the command line before the script (sbatch --time=1:00:00 train.sbatch); for the same option, the later one wins.
  2. A run named after --job-name (or the script's file name: train.sbatch → train) is created with 16 GPUs on 2 machines in the pool h100, 8 cores and 64 GiB per GPU (the per-node 64 cores and 512 GiB divided by 8 GPUs), and a 2-day limit.
  3. The script is mounted read-only at /astra/job.sbatch, and each worker runs bash /astra/job.sbatch followed by any arguments you gave after the script's name.
  4. Each worker runs the whole script once on its machine. That is what srun with one task per node would do, so drop srun from the script's commands: there is no srun inside the container.

Options Astraeus does not use are named on standard error and left out: astra: not used on Astraeus: --mail-type=END --dependency=afterok:41.

Run one command: srun#

srun submits a command as a run, then follows its first worker's log until it ends, and exits non-zero if it did not complete:

$ ASTRA_IMAGE=nvcr.io/nvidia/cuda:12.6.0-base-ubuntu24.04 astra slurm srun -N 1 --gpus=1 nvidia-smi -L
srun: job nvidia-smi-48213 queued
GPU 0: NVIDIA H100 80GB HBM3 (UUID: GPU-5c1f…)

nvidia-smi-48213-0: Completed — Exited with code 0

The run is named --job-name, or <command>-<process id>. Options come first; the first word that is not an option starts the command, and everything after it belongs to the command.

See and cancel jobs#

$ astra slurm squeue
             JOBID  PARTITION     USER ST    CLUSTER  REASON
             llama          -        -  R   gpu-east  2 of 2 workers running
              eval          -        - PD   gpu-east  Waits for namespace quota: gpus 16 in use + 8 asked > 16
$ astra slurm scancel eval
$ astra slurm sinfo
PARTITION    AVAIL  NODES STATE    NODELIST
h100            up      1 drain    gpu-10
h100            up      3 idle     gpu-07,gpu-08,gpu-09
Command What it shows or does
squeue The workspace's runs that have not ended: JOBID is the run's name, ST its state (PD Pending, CF Starting, R Running), REASON its reason. PARTITION and USER are always -.
scancel <job>… Deletes the runs: their workers are stopped, and the runs and their logs are removed.
sinfo The machines the workspace may use, by pool: idle is up (whether or not it is busy), drain is cordoned, down is down.

scancel deletes

Unlike Slurm, there is no record of a cancelled job afterwards: the run and its logs are gone.

#SBATCH options#

Option Becomes
-J, --job-name The run's name.
-N, --nodes=N machines: N (one worker per machine) when N > 1, or N = 1 with GPUs.
-N, --nodes=N-M A range: topology.min_nodes: N, max_nodes: M.
--gres=gpu:G, --gres=gpu:<type>:G G GPUs per node; the type is not used for selection. Other --gres are named as not used.
--gpus-per-node=G, --gpus-per-task=G G GPUs per node.
-G, --gpus=G (or <type>:G) G GPUs in total.
-c, --cpus-per-task=C C cores per worker; with GPUs, divided by the GPUs per machine (at least 1 per GPU).
--mem=M Memory per worker (64G, 512M, 1.5T; a plain number is MiB); with GPUs, divided by the GPUs per machine.
--mem-per-cpu=M M × cores, when --mem is not given.
-t, --time time_limit_seconds. Minutes (30), MM:SS, HH:MM:SS, D-HH:MM:SS, or 1h30m.
-a, --array=0-99%10 array: {size: 100, max_parallel: 10}. See Arrays.
-p, --partition=P Only machines labelled pool=P.
-w, --nodelist=a,b Only these machines (comma-separated names; no gpu-[1-4] ranges).
--exclusive At least min(8, GPUs) GPUs on each machine.
--nice=N priority: −N.
--container-image=I The image.
-n, --ntasks, --ntasks-per-node, -o, --output, -e, --error, -A, --account, -q, --qos, --export, -D, --chdir Ignored silently: Astraeus runs one worker per machine, output goes to the run's log, and accounts are workspaces.
Anything else (--dependency, --mail-type, --constraint, --reservation…) Named as not used, and ignored.

Every run made by sbatch or srun has restart_policy: Never and no start policy: its workers start independently, each as soon as it fits, and a failure fails the run.

Multi-node scripts are not started as a gang

Slurm allocates all nodes before a job starts. sbatch runs on Astraeus start each worker when it fits, so on a busy cluster one machine's worker may start well before the other's. Frameworks that wait for every peer (torchrun) then wait, up to their rendezvous timeout. For a gang, submit a run with start: Gang (the console does this for multi-machine runs); see Multi-machine runs.

Arrays#

--array sets the array's size from the range: 0-99 and 1-100 are both 100 copies, 3 is 4 copies (0-3), 1,5,9 is 3 copies; %N caps how many run at once. Copies are always numbered 0 to size − 1, in ASTRAEUS_ARRAY_TASK_ID; SLURM_ARRAY_TASK_ID is not set. A script using a range that does not start at 0, or a list, must map the index itself:

INDICES=(1 5 9)
i=${INDICES[$ASTRAEUS_ARRAY_TASK_ID]}

Variables inside a worker#

Every worker of every run — not only those made by sbatch — gets these:

Slurm Astraeus
SLURM_JOB_ID, SLURM_JOBID, SLURM_JOB_NAME The run's name.
SLURM_PROCID, SLURM_NODEID The worker's rank (RANK).
SLURM_LOCALID 0: one worker per machine.
SLURM_NTASKS, SLURM_NPROCS, SLURM_NNODES, SLURM_JOB_NUM_NODES The number of workers (WORLD_SIZE).
SLURM_NTASKS_PER_NODE 1.
SLURM_CPUS_PER_TASK The worker's cores.
SLURM_MEM_PER_NODE The worker's memory, MiB.
SLURM_GPUS_ON_NODE, SLURM_GPUS_PER_TASK The worker's GPUs.
SLURMD_NODENAME The machine's name.
SLURM_CLUSTER_NAME astraeus.
— MASTER_ADDR, MASTER_PORT (29500): the leader's machine address.

Not set: SLURM_NODELIST, SLURM_JOB_NODELIST, SLURM_ARRAY_TASK_ID, SLURM_ARRAY_JOB_ID, SLURM_STEP_*, SLURM_SUBMIT_DIR. Libraries that derive the leader from SLURM_NODELIST need MASTER_ADDR instead. See Inside a worker.

What is not supported#

  • Job dependencies (--dependency), heterogeneous jobs, job steps, sacct, salloc, sbcast, sattach.
  • Several tasks per node (--ntasks-per-node > 1): run your own launcher (torchrun --nproc-per-node) inside the one worker.
  • Mail, accounts, QoS, licences: workspaces, quotas and priorities take their place.
  • Mounting drives or credentials from #SBATCH: use a run specification (Submit a run).

Troubleshooting#

Symptom Cause Fix
astra: the container image: #SBATCH --container-image=… (as with Pyxis), or ASTRA_IMAGE No image. Add #SBATCH --container-image=… or set ASTRA_IMAGE.
astra: the batch script (a file) is missing No readable script among the arguments. Give the script's path.
srun: command not found in the log The script calls srun, which is not in the image. Remove srun: each worker already runs the script on its machine.
Every array copy does the same work The script reads SLURM_ARRAY_TASK_ID. Read ASTRAEUS_ARRAY_TASK_ID.
--nodelist=gpu-[1-4] matches nothing Ranges are not expanded. List the names: gpu-1,gpu-2,gpu-3,gpu-4.
The job never finds its peers Workers of a multi-node script started at different times, or MASTER_ADDR does not reach the leader. Submit as a gang (Multi-machine runs).