Priorities, preemption and checkpoints#
When there is more work than GPUs, the queue decides who runs next. Priority puts urgent work first and may stop lower-priority work to make room; fair share keeps one workspace from taking everything; backfill lets short runs use room that a large run is waiting for. This page explains each, and how to write a run that survives being stopped: it saves a checkpoint when asked and resumes from it.
How the queue orders runs#
The scheduler looks at every waiting run at least every 15 seconds, and whenever something changes. Runs are ordered by:
- Priority, higher first.
- Fair share: the workspace using the smallest part of its quota on that cluster first, divided by its weight.
- Submission time, earlier first.
The first run in that order that cannot start now, but could on this cluster, is the head. The head is protected: nothing behind it may take the room it waits for, so a stream of small runs cannot starve a large one. The head says so: Next to run (priority 0): starts by 2026-10-01 14:00 UTC as running work ends. Runs behind it say Queued behind <run>….
A run over its workspace's quota waits (Waits for namespace quota: gpus 6 in use + 4 asked > 8) without blocking anyone, and is never the head.
Set a priority#
priority is an integer from −1000 to 1000; the default is 0. A workspace may ask for at most its maximum priority on that cluster, which an organisation admin sets on the workspace's access to the cluster (see Workspaces, quotas and pools). The default maximum is 0, so without a grant you can lower a run's priority but not raise it. A higher value is refused: 403 PRIORITY_NOT_ALLOWED: priority 50 is above namespace …'s maximum of 10.
New run → More options → Priority. The hint says the maximum: Higher runs first and may preempt lower. Up to 10 here. The panel on the right repeats it as Max priority.
astra astraeus run has no priority flag; astra slurm sbatch --nice=<n> sets priority −n.
"priority": 100 in spec.
The same cap applies to the run templates of schedules and replica groups.
Preemption#
When the head has a higher priority than work that is running, the scheduler stops the least it can of that work to make room:
- Who: only running work of strictly lower priority, from any workspace, on machines the head may use. Equal priority never preempts.
- Units: a gang is stopped whole; an independent worker alone.
- Order: lowest priority first; among equals, the most recently started (the least work lost). Units that turn out not to be needed are kept.
- How: each victim moves to
Stoppingwith the reasonPreempted by higher-priority work (<run>). Its machine sendsSIGTERM, waits the stop grace (30 s by default), then sendsSIGKILL. - Afterwards: the victim is
Cancelled, then queued again (Preempted by <run>; queued again) without spending its restart budget. A gang is queued again only once none of its workers is still stopping. Its run showsPreempted by higher-priority work (<run>); waiting to run again. - The room is held: while victims stop, the head says
Preempting 2 lower-priority worker(s) of <run> to start, thenStarting: waiting for 2 preempted worker(s) to stop, and nothing else may take that room.
Another workspace's run is never named to you: reasons say a job in another namespace instead.
Stop grace
The 30-second grace is a setting of each machine (ASTRAEUS_STOP_GRACE in the machine's agent configuration), not of the run. It applies to every stop: preemption, time limit, delete.
Backfill#
Backfill lets runs behind the head start now when they cannot delay it. The scheduler computes when the head will start at the latest — its shadow time — from the time limits of the work that is running, and holds the room the head will take then. A run behind it may start now if:
- it fits in what the head does not need, or
- its own time limit ends it before the shadow time.
A run with no time limit cannot be known to end in time, so it only gets room the head does not need. The promise is per resource, not per machine: a CPU-only run may use a machine whose GPUs are promised.
So, to start sooner on a busy cluster, give runs an honest time_limit_seconds. If the running work has no time limits, the head's start time is unknown: Next to run (priority 0): waits for running work to end, and only room it does not need is used.
Fair share and quotas#
Each workspace has, on each cluster, a quota (GPUs, CPU cores, memory, workers at once) and a weight. Among runs of equal priority, the workspace whose dominant share — the largest fraction of any quota it is using — divided by its weight is smallest goes first. A workspace without a quota has a share of 0.
Quota is a hard ceiling: a run that would exceed it waits, and records a Quota reached: … entry (and event) the first time. A run that alone exceeds the quota says Waits for more than the namespace's quota allows at all (…): raise the quota or ask for less. See Workspaces, quotas and pools.
Write a checkpoint-friendly run#
A run can be stopped at any time: preempted, past its time limit, its machine lost, or restarted by a gang failure. Write it so that stopping costs little:
- Save checkpoints to a drive, not the container's filesystem, which is lost when the worker stops. Mount a drive read-write.
- Resume from the latest checkpoint at start, if there is one.
- Save on
SIGTERMwithin the stop grace (30 s by default), then exit. Also save periodically: a lost machine sends no signal. - Write atomically: write to a temporary file and rename it, so a kill mid-write leaves the previous checkpoint intact.
- Make sure the signal reaches your program. With
"command": "bash", "args": ["-c", "…"], start your program withexecso it replaces the shell and receivesSIGTERMitself.torchrunpassesSIGTERMon to its processes.
import os, signal, sys, threading
import torch
CKPT_DIR = "/ckpt/" + os.environ.get("SLURM_JOB_NAME", "run") # the run's name
CKPT = os.path.join(CKPT_DIR, "latest.pt")
stop = threading.Event()
signal.signal(signal.SIGTERM, lambda signum, frame: stop.set())
def save(step, model, opt):
os.makedirs(CKPT_DIR, exist_ok=True)
tmp = CKPT + ".tmp"
torch.save({"step": step, "model": model.state_dict(), "opt": opt.state_dict()}, tmp)
os.replace(tmp, CKPT) # atomic: a kill mid-save keeps the previous one
model, opt = build_model(), build_optimizer()
start = 0
if os.path.exists(CKPT):
state = torch.load(CKPT, map_location="cpu")
model.load_state_dict(state["model"])
opt.load_state_dict(state["opt"])
start = state["step"] + 1
for step in range(start, TOTAL_STEPS):
train_step(model, opt, step)
if stop.is_set():
save(step, model, opt)
sys.exit(0)
if step % 500 == 0:
save(step, model, opt)
In a multi-machine run, have rank 0 write the checkpoint (or use your framework's distributed checkpointing), and keep the save shorter than the grace.
A preemptible fine-tune#
A fine-tune at low priority that uses idle GPUs, gives way to urgent work, and picks up where it stopped:
- New run: Name
ft-lowprio, Image, Commandbash, then Edit as JSON to setargsas in the API tab. - GPUs
8, Drivescheckpoints:/ckpt. - More options: Priority
-100, Time limit48h. - Start run.
with #SBATCH --container-image=… and #SBATCH --gres=gpu:8 in finetune.sbatch (see Slurm compatibility). astra cannot mount the drive: use a template that has it, or the API.
{
"metadata": { "name": "ft-lowprio" },
"spec": {
"priority": -100,
"task_template": {
"image": "nvcr.io/nvidia/pytorch:24.08-py3",
"command": "bash",
"args": ["-c", "exec torchrun --nproc-per-node=$ASTRAEUS_GPU_COUNT /ckpt/code/train.py"],
"time_limit_seconds": 172800,
"datavolume_refs": [{ "name": "checkpoints", "mount_path": "/ckpt", "mode": "ReadWrite" }],
"requested_resources": {
"gpu_requests": { "count": 8, "min_per_machine": 8 },
"per_gpu": { "cpu_cores": 8, "memory_bytes": 68719476736 }
}
}
}
}
When an urgent run preempts it, the run's page shows Paused to make room for higher-priority work; it is queued again and starts from /ckpt/ft-lowprio/latest.pt when room returns. The preemption does not count against its 10 restarts.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
403 PRIORITY_NOT_ALLOWED |
Above the workspace's maximum priority. | Lower it, or ask an organisation admin to raise the maximum. |
A small run waits Queued behind <big run>, which starts by … while GPUs look free |
The free GPUs are held for the head. | Give the small run a time limit that ends before that time. |
| A high-priority run does not preempt | The running work has equal or higher priority, or stopping it would not make room. | Check the per-machine reasons; a gang needs room for all its workers. |
| Work restarts from the beginning after a preemption | It does not resume from a checkpoint, or the save did not finish within the grace. | Follow Write a checkpoint-friendly run. |
SIGTERM never reaches the program |
A shell is PID 1 and does not forward it. | exec your program from the shell. |