Share a cluster among teams#
At data-centre scale, sharing is rarely just two teams on one kind of machine. You have several hardware classes — H100s for large training, A100s for everything else, L40S for inference — and several teams, each with its own quota, fair-share weight and ceiling on priority. And sometimes a team needs a whole rack held for a launch date, not just a bigger quota. This recipe pools a fleet by hardware class, grants it to three teams, and reserves machines for a deadline.
For the mechanics of workspaces, quotas, fair share and preemption between two teams on one pool, read Share a GPU fleet between teams first; this recipe builds on it rather than repeating it.
Before you begin#
- An organisation owner or admin: workspaces, cluster access terms, machine labels and reservations are theirs to set.
- A cluster with more than one kind of machine. The example: 6 machines
with 8 × H100 (pool
h100, 48 GPUs), 4 with 8 × A100 (poola100, 32 GPUs), and 2 with 8 × L40S (pooll40s, 16 GPUs). - For the API: an API token,
curlandjq(How the recipes are written).<org>and<cluster>stand for yours.
1. Pool the fleet by hardware class#
$ for m in h100-01 h100-02 h100-03 h100-04 h100-05 h100-06; do
curl -sS -X PUT "https://api.astralyx.cloud/v1/organizations/<org>/clusters/<cluster>/machines/$m/placement" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' -d '{"pool": "h100"}' >/dev/null
done
$ for m in a100-01 a100-02 a100-03 a100-04; do … -d '{"pool": "a100"}' …; done
$ for m in l40s-01 l40s-02; do … -d '{"pool": "l40s"}' …; done
Or label each machine from its page, under Where it is, as in Pools, labels and topology. Running work stays where it is; the label only changes where new work can land.
2. Give each team its terms#
Vision trains large models on H100s only. Research uses A100s for everything, and may also use H100s when Vision is not using all of them. Eval runs short, deadline-bound jobs on any machine and may jump the queue.
| Term | vision |
research |
eval |
|---|---|---|---|
Pools (node_selector) |
pool=h100 |
none: any machine | none: any machine |
| Quota (GPUs) | 40 | 48 | 16, 32 workers |
| Weight | 2 | 1 | 1 |
| Max priority | 10 | 10 | 60 |
Research is given no pool so it can fall back to idle H100 capacity; its quota (48) still caps what it can hold at once. A workspace whose access lists no pools may use every machine of the cluster, pools included (see Pools, labels and topology).
$ ORG=https://api.astralyx.cloud/v1/organizations/<org>
$ curl -sS -X POST "$ORG/workspaces/vision/clusters" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d '{"cluster": "<cluster>", "quota": {"gpus": 40}, "weight": 2, "node_selector": {"pool": "h100"}, "max_priority": 10}'
$ curl -sS -X POST "$ORG/workspaces/research/clusters" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d '{"cluster": "<cluster>", "quota": {"gpus": 48}, "weight": 1, "max_priority": 10}'
$ curl -sS -X POST "$ORG/workspaces/eval/clusters" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d '{"cluster": "<cluster>", "quota": {"gpus": 16, "tasks": 32}, "weight": 1, "max_priority": 60}'
See Fair share and quotas for how the dominant share, divided by weight, orders equal-priority work across all three, and Give each workspace access to the cluster, on its terms for every field.
3. Reserve a rack for a launch date#
Vision has a model that must finish training before a launch on the 10th. Reserve four H100 machines for Vision from the 5th at 18:00 to the 10th at 08:00 UTC, so nothing but Vision's work lands there once the window opens:
- Clusters → the cluster → Reservations → Reserve machines.
- Or a pool:
pool=h100, or tickh100-01…h100-04by name. - From 2026-10-05 18:00, To 2026-10-10 08:00 (your browser's time zone).
- For: tick
Vision. Why:launch training. - Reserve.
$ curl -sS -X POST "$ORG/clusters/<cluster>/reservations" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"machines": ["h100-01", "h100-02", "h100-03", "h100-04"],
"start": "2026-10-05T18:00:00Z", "end": "2026-10-10T08:00:00Z",
"workspaces": ["vision"], "reason": "launch training"}' | jq -r .metadata.name
r-3f9a0c1d
Before the window, Research and Eval may still use those four machines — but only with runs whose time limit ends before 18:00 on the 5th; a run with no time limit is kept off them as soon as the reservation is ahead of it. A reservation decides placement, not eviction: it does not stop work already running when the window opens. During the window, only Vision's runs are placed there; after it, anyone's again.
A reservation does not move Vision's runs onto those machines by itself — its runs still go wherever they fit. To use exactly the reserved machines:
"requested_resources": {
"gpu_requests": {"count": 32, "min_per_machine": 8},
"machine_selection": {"mode": "Any", "names": ["h100-01", "h100-02", "h100-03", "h100-04"]}
}
Submit it any time; it waits Pending until the window opens and the
machines are free. See Reservations.
4. Hold machines for maintenance#
A reservation with no workspaces holds machines for nobody — the pattern for a firmware upgrade or a planned power-feed test:
$ curl -sS -X POST "$ORG/clusters/<cluster>/reservations" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"pool": {"pool": "a100"}, "start": "2026-10-12T02:00:00Z", "end": "2026-10-12T06:00:00Z", "workspaces": [], "reason": "firmware upgrade"}'
From the moment it exists, no new run without a time limit lands on the
a100 pool; when the window opens, cordon the machines if anything is
still running there (a reservation never evicts). See
Plan maintenance and
Update, drain and remove.
5. Watch usage across teams and pools#
Organisation → Usage shows GPU-hours and cost per workspace over time; each workspace's Compute → GPUs shows, GPU by GPU, who holds what — another team's work shown only as taken capacity, never its name. See Usage, cost and budgets.
6. Change terms later#
$ curl -sS -X PATCH "$ORG/workspaces/research/clusters/<cluster>" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d '{"quota": {"gpus": 56}, "weight": 1, "node_selector": {}, "max_priority": 10}'
PATCH replaces only the terms you send, each whole. New terms apply to
what is placed next; runs already running are not stopped because a quota
went down. PUT/DELETE on a reservation works the same way — the
reservation's name stays.
What you get#
- Three teams, three ceilings on priority, and fair share that keeps one team from starving another, across three hardware classes.
- A guaranteed rack for a launch date, without evicting anyone who was using it before the window, and without Research or Eval ever seeing Vision's run by name.
- A maintenance window that holds machines empty of its own accord once a reservation exists.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
| A workspace sees no machines | Its pools match no machine's labels. | Check the labels under Where it is, or the workspace's Pools. |
| The reserved run does not start when the window opens | Work started before the window is still running there. | Wait for it, or stop it; reservations do not evict. |
| Runs avoid the reserved machines days ahead of the window | They have no time limit, so nothing says they would end in time. | Give them time limits, or reserve closer to the date. |
400 INVALID_RESERVATION: name machines, or a pool |
Neither machines nor pool was given. |
Give one. |
404 WORKSPACE_NOT_FOUND creating a reservation |
A slug in workspaces is not the organisation's. |
Use the workspace's short name. |
403 PRIORITY_NOT_ALLOWED |
A run asked for a priority above its workspace's maximum. | Lower it, or raise the workspace's Max priority. |