Skip to content

Runs and workers#

A run is a piece of work you ask Astraeus to do: an image, a command, and what it needs — "train this on 16 GPUs for at most 12 hours". The scheduler turns it into workers, one container each on one machine, and places them. A worker's index in its run is its rank. In the API a run is a job (/jobs, also /runs) and a worker is a task (/tasks, also /workers).

Use a run when:

  • You train or fine-tune a model, on one GPU or across many machines (distributed training starts all workers together, as a gang).
  • You process a batch: preprocessing, evaluation, batch inference — work that ends.
  • You sweep parameters: an array runs the same template N times, each with its index, at most M at once.
  • You keep something running: a notebook server, a development container, an inference server — a run with lifetime Service is restarted whenever it ends, until you delete it.
  • Your work has parts: a parameter server and trainers, an engine and its clients — worker groups of different images and sizes in one run, one starting after another.

Workers and ranks#

Workers are named <run>-<rank>: train-0, train-1. In a run with worker groups, <run>-<group>-<rank>. Every worker gets its rank and the run's size in its environment (RANK, WORLD_SIZE, and Slurm's SLURM_PROCID, SLURM_NTASKS…), and, in a multi-worker run, the leader's address (MASTER_ADDR, MASTER_PORT = 29500) — what torchrun and most launchers read. See Inside a worker.

The shape: GPUs, not machines#

A run asks for GPUs for the whole run (gpu_requests.count). Unless you say otherwise, the scheduler picks the shape: how many machines, and how many GPUs on each — the fewest machines on the best network that has room, tried in this order: one machine, one NVLink domain, one InfiniBand fabric (and rack), one RDMA network, one rack, Ethernet last. CPU cores and memory can follow the GPUs (per_gpu), so the total fits whatever shape is chosen.

You can constrain it instead:

You want Ask for
Exactly N machines, GPUs split evenly machines: N
Never fewer than K GPUs on one machine gpu_requests.min_per_machine: K
A GPU model, a minimum GPU memory gpu_requests.models, gpu_requests.min_memory_gb
Only healthy GPUs gpu_requests.healthy_only: true
Never slower than InfiniBand between workers network.interconnect: "infiniband"
The whole run inside one rack, fabric or NVLink domain topology.keep_within

See GPUs and placement and Multi-machine runs.

How workers start: start#

Value Behaviour Default for
Gang All or nothing: every worker placed together, and started together behind a barrier. A gang is Starting until every worker runs; it waits up to 15 minutes at the barrier. Runs that may span several machines
MinAvailable Starts once min_available workers can start together; the rest join as room appears. For Spark and elastic training. —
Independent Each worker starts when it fits. Runs on one machine, arrays

When a worker fails#

Two settings decide.

restart_policy on the worker template — whether a failed worker is retried at all:

Value Behaviour
OnFailure (default) A failed or lost worker is restarted, up to 10 times, waiting 10 s before the first restart and doubling up to 5 min. A worker that runs 10 minutes has its count reset.
Never A failed worker is not restarted. The CLIs set this.
Always Also restarted when it completes.

on_failure on the run — what a restart covers:

Value Behaviour Default for
RestartJob Restart every worker: synchronous training does not survive the loss of a member. Gangs
RestartTask Restart only the worker that failed. In a gang, losing the leader still restarts all. Everything else
FailJob No restart: the run fails and its other workers are stopped. —

A worker stopped by its time limit is never restarted. A worker preempted for higher-priority work goes back to the queue without spending its restart budget. See Priorities, preemption and checkpoints.

How long: lifetime#

Value Behaviour
Batch (default) The run ends when its workers end: Completed when all complete, Failed under its failure policy.
Service Runs until you delete it; restarted whenever it ends. Never Completed.

Run states#

State Meaning
Pending No worker runs yet: waiting in the queue, or being prepared.
Starting A gang placed whole; its workers are preparing and start together.
Running At least one worker runs (every one, for a gang).
Completed Every worker completed successfully.
Failed The run failed under its policy.
Cancelled Cancelled by a user.

Workers have more states (Preparing, Pulling, ReadyToStart, Stale, Down…). Every state comes with a reason in words. See Run and worker states.

The Runs page of a workspace, with running, pending and failed runs

Create one#

A distributed PyTorch training run on 16 GPUs, letting the scheduler pick the machines, with 12 CPU cores and 120 GiB of memory per GPU and a 12-hour limit:

  1. New → Run.
  2. Name llama-ft, Image nvcr.io/nvidia/pytorch:24.08-py3.
  3. GPUs 16, Machines Automatic, CPU cores per GPU 12, Memory per GPU 120Gi.
  4. Under More options: Time limit 12h. Start stays Automatic: a gang, because the run may span machines.
  5. Choose Edit as JSON (worker groups, drives, probes…) and set command and args as in the API tab. Arguments are passed as they are: no shell expands $RANK unless the command is sh -c.
  6. Start run.
$ astra astraeus run --name llama-ft \
    --image nvcr.io/nvidia/pytorch:24.08-py3 \
    --gpus 16 --cpus 12 --mem 120G \
    --time 12h \
    -- sh -c 'torchrun --nnodes=$WORLD_SIZE --node-rank=$RANK --nproc-per-node=gpu \
              --master-addr=$MASTER_ADDR --master-port=$MASTER_PORT train.py'
llama-ft submitted to default: https://console.astralyx.cloud/o/acme/w/default/jobs/default/llama-ft

With GPUs, --cpus and --mem are per GPU. astra astraeus run starts its workers independently, not as a gang; for a gang, use the API tab's specification (or save it as a template and pass --template).

run.json
{
  "metadata": {"name": "llama-ft"},
  "spec": {
    "start": "Gang",
    "on_failure": "RestartJob",
    "task_template": {
      "image": "nvcr.io/nvidia/pytorch:24.08-py3",
      "command": "sh",
      "args": ["-c", "torchrun --nnodes=$WORLD_SIZE --node-rank=$RANK --nproc-per-node=gpu --master-addr=$MASTER_ADDR --master-port=$MASTER_PORT train.py"],
      "requested_resources": {
        "gpu_requests": {"count": 16},
        "per_gpu": {"cpu_cores": 12, "memory_bytes": 128849018880}
      },
      "time_limit_seconds": 43200
    }
  }
}
$ curl -sS -X POST "$WS_API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @run.json

$WS_API is your workspace's view of the cluster, https://console.astralyx.cloud/api/v1/orgs/<org>/workspaces/<workspace>/clusters/<cluster>/api.

train.py must be in the image (or on a drive). For a complete, step-by-step example, see Your first run.