The queue#
When more work is waiting than the machines can hold, the scheduler's queue decides what starts next. It is one queue for the whole cluster, in the style of an HPC scheduler: priorities, fair share between workspaces, quotas, all-or-nothing starts for distributed runs, backfill into gaps, and preemption of lower-priority work. This page explains the model; the settings are in Priorities, preemption and checkpoints and Workspaces, quotas and pools.
Use the queue's controls when:
- Urgent work must start now: give it a higher
priority; it may stop lower-priority work to make room. - Teams compete for the same GPUs: give each workspace a quota and a weight; waiting runs of the least-served workspace go first.
- A large run keeps waiting behind small ones: it doesn't — the queue
holds room for it. Give small runs a
time_limit_secondsso they can still use the gaps. - A distributed run must start whole or not at all: it does by default (a gang); no GPUs sit idle half-allocated.
- You need machines free at a known time: an admin reserves them for a window.
The order#
Every waiting run is ordered by, in turn:
- Priority, highest first. A run's
priority(default0) can be at most its workspace's maximum priority on that cluster; asking more is refused with403 PRIORITY_NOT_ALLOWED. - Fair share: among equal priorities, the workspace least served first. A workspace's share is the largest fraction of its quota it holds now — GPUs, cores, memory or workers, whichever is highest — divided by its weight. A workspace with no quota counts as holding none of it.
- Submission time, earliest first.
Quotas#
A workspace's quota on a cluster caps what it holds at once. A run that would go over it waits without blocking anyone: runs behind it, from the same workspace or others, may start. Its waiting workers' reason says so, for example Gang waits for namespace quota: gpus 12 in use + 8 asked > 16, and the run's history records Quota reached. When a run asks for more than the quota allows even with nothing else running, it says more than the namespace's quota allows at all and waits until the quota is raised.
The quota is checked again when a worker is bound to a machine, in the same transaction: two placements on different machines can never together exceed it.
The head of the queue#
The first run in that order that cannot start now becomes the head. Nothing else may take the room it is waiting for. Two mechanisms make room:
Preemption. If lower-priority work is running, the scheduler stops the fewest units of it whose removal lets the head start — a gang whole, an independent worker alone — lowest priority first, and among equals the most recently started (the least work lost). Stopped workers wind down gracefully: the machine sends the stop signal and waits its stop grace (30 s by default) before killing them, time to save a checkpoint. They go back to the queue with the reason Preempted by higher-priority work and do not spend their restart budget.
Backfill. Otherwise the head is promised the room that running work frees earliest, computed from the running work's time limits: its shadow time. Runs behind the head may still start now if they fit around that promise, or if their own time limit makes them end before it. A run with no time limit cannot be known to end in time, so it only gets room the head does not need.
Tip
Set time_limit_seconds on every run that has a known bound. It is what
lets the queue start it early, in the gaps before large runs.
Gangs#
A run that may span several machines starts as a gang unless you say otherwise: all its workers are placed together or not at all, then start together behind a barrier. The scheduler never holds half a gang's GPUs while it waits for the rest. While a gang waits, its reason says Waiting for room for all its workers at once. See Runs and workers.
Reservations#
An organisation admin can reserve machines for a window — for some workspaces' large run, or for maintenance (reserved for no one). Before the window starts, only work certain to end in time (by its time limit) is placed on those machines. See Reservations.
Why a run waits#
Every waiting run and worker carries a reason in words. The console's run page explains it and, under Why it is not running yet, lists what each machine lacks.

| Reason (as shown in the console) | What it means |
|---|---|
| Waiting for a machine with room | Enough GPUs, cores and memory are not free yet. |
| Waiting for room for all its workers at once | A gang waits until every worker can be placed. |
| The workspace has used its quota on this cluster | Over quota; it starts when other work of the workspace ends or the quota is raised. |
| Paused to make room for higher-priority work | Preempted; it runs again when there is room. |
| Waiting for a better network | It asked to wait up to a time for a tighter network than there is room on now. |
| No machine has … GPUs | No machine the workspace may use has the GPU model asked for. |
See Troubleshooting and FAQ for every reason and its fix.
Try it#
Start a long run, then a higher-priority one that needs its GPUs. Your
workspace's maximum priority must be at least 50 (its Settings show
it).
- New → Run: Name
background, Imagebusybox:1.36, Commandsleep, Arguments3600, GPUs: every GPU of one machine, CPU cores per GPU1, Memory per GPU1Gi. Start run. - New → Run again: Name
urgent, same image, Commandsleep, Arguments60, the same GPUs, cores and memory. Under More options, Priority50and Time limit10m. Start run. - Open
background: its worker is stopped with Paused to make room for higher-priority work, and queued again.urgentruns.
astra astraeus run has no priority flag; Slurm's --nice=-50 sets
priority 50. Use every GPU of one machine (8 here):
$ astra astraeus run --name background --image busybox:1.36 --gpus 8 --cpus 1 --mem 1G -- sleep 3600
$ astra slurm sbatch -J urgent --container-image=busybox:1.36 -G 8 -c 8 --mem 8G --nice=-50 -t 10 urgent.sh
Submitted batch job urgent
$ astra astraeus runs
with urgent.sh holding sleep 60. astra astraeus runs shows
urgent running and background waiting again. Its worker's reason
(astra astraeus workers background) is Preempted by urgent; queued
again.