Skip to content

Your first run#

This walkthrough runs a small PyTorch training loop on one GPU. You write the script, describe the run in a spec file, submit it, follow its state and log, delete it, and start it again. Every field used is explained; the full list is in the Run specification.

Before you begin#

  • A workspace with at least one machine with an NVIDIA GPU that is Up under Astraeus → Machines. See the Quick start. No GPU? Set the GPU count to 0 below: the script falls back to the CPU.
  • For the CLI tab: astra, signed in and working in your workspace. See Install the CLI.
  • For the API tab: a personal API token (account menu → Account & tokens → API tokens → New token) in ASTRAEUS_TOKEN.
  • jq, to build the spec file below.

The image is large

nvcr.io/nvidia/pytorch:24.08-py3 is several gigabytes. The first run on a machine spends minutes in Pulling; later runs on that machine start at once.

1. Write the script#

The script trains a small network on random data, so it needs no dataset. It prints its rank, its machine and its GPU, then its progress every 200 steps.

train.py
import os, time, torch

device = "cuda" if torch.cuda.is_available() else "cpu"
gpu = torch.cuda.get_device_name(0) if device == "cuda" else "no GPU"
print(f"rank {os.environ['RANK']} of {os.environ['WORLD_SIZE']} on {os.environ['ASTRAEUS_NODE_NAME']}: {gpu}", flush=True)

steps = int(os.environ.get("STEPS", "2000"))
model = torch.nn.Sequential(torch.nn.Linear(4096, 4096), torch.nn.ReLU(), torch.nn.Linear(4096, 10)).to(device)
opt = torch.optim.AdamW(model.parameters(), lr=1e-3)
start = time.time()
for step in range(1, steps + 1):
    x = torch.randn(512, 4096, device=device)
    y = torch.randint(0, 10, (512,), device=device)
    loss = torch.nn.functional.cross_entropy(model(x), y)
    opt.zero_grad()
    loss.backward()
    opt.step()
    if step % 200 == 0:
        print(f"step {step}/{steps} loss {loss.item():.4f} {step / (time.time() - start):.0f} it/s", flush=True)
print("done", flush=True)

Every worker gets RANK, WORLD_SIZE and ASTRAEUS_NODE_NAME in its environment, among others. See Inside a worker.

2. Describe the run#

A run is a JSON object: metadata.name and a spec. The script travels inside the spec as a file mounted into the container, so the run needs nothing on the machine. Build the spec from the script with jq:

$ jq -n --rawfile code train.py '{
    metadata: {name: "first-train"},
    spec: {
      task_template: {
        image: "nvcr.io/nvidia/pytorch:24.08-py3",
        command: "python",
        args: ["/workspace/train.py"],
        env: {STEPS: "3000"},
        configs: [{mounts: ["/workspace/train.py"], value: $code}],
        requested_resources: {
          gpu_requests: {count: 1},
          per_gpu: {cpu_cores: 8, memory_bytes: 34359738368}
        },
        restart_policy: "Never",
        time_limit_seconds: 3600
      }
    }
  }' > run.json

The result, with the script shortened:

run.json
{
  "metadata": {"name": "first-train"},
  "spec": {
    "task_template": {
      "image": "nvcr.io/nvidia/pytorch:24.08-py3",
      "command": "python",
      "args": ["/workspace/train.py"],
      "env": {"STEPS": "3000"},
      "configs": [{"mounts": ["/workspace/train.py"], "value": "import os, time, torch\n…"}],
      "requested_resources": {
        "gpu_requests": {"count": 1},
        "per_gpu": {"cpu_cores": 8, "memory_bytes": 34359738368}
      },
      "restart_policy": "Never",
      "time_limit_seconds": 3600
    }
  }
}

The fields#

Field Value here What it does
metadata.name first-train The run's name: lower-case letters, digits and dashes, unique in the workspace. Its workers are named <name>-<rank>: here first-train-0.
spec.task_template The worker template: what each worker runs and needs. A run without worker groups has one template.
image nvcr.io/nvidia/pytorch:24.08-py3 The container image. A private registry needs a registry credential (registry_secret_ref).
command python Replaces the image's entrypoint. Empty: the image's own.
args ["/workspace/train.py"] The command's arguments, one string each. No shell is involved: to use pipes or $VARS, run sh -c.
env {"STEPS": "3000"} Environment variables, as strings. Variables Astraeus sets (RANK, MASTER_ADDR…) are only set when you do not set them yourself.
configs One file Literal files written into the container at each path in mounts. For small files such as scripts and configuration — the whole specification is kept by Astraeus and limited in size, so keep data on a drive.
requested_resources.gpu_requests.count 1 GPUs for the whole run. Astraeus decides how many machines they span unless you say. 0: a CPU-only run.
requested_resources.per_gpu 8 cores, 32 GiB CPU cores and memory for each GPU, so the total follows the shape Astraeus picks. Without GPUs, use cpu_cores and memory_bytes on requested_resources instead. Memory is in bytes: 34359738368 is 32 GiB.
restart_policy Never A worker that fails is not restarted, and the run fails. The default, OnFailure, restarts a failed worker up to 10 times, waiting 10 s and doubling up to 5 min between attempts.
time_limit_seconds 3600 The worker is stopped after one hour and is not restarted. A time limit also lets the queue start your run sooner, in the gaps before large runs (backfill).

Fields not set take their defaults: start is Independent for a run that fits on one machine, priority is 0, lifetime is Batch (the run ends when its workers end). A GPU worker's /dev/shm is half its memory, between 1 GiB and 64 GiB — what PyTorch's data loaders need.

Unknown fields are refused with 400 UNKNOWN_FIELD, so a typo never passes silently.

3. Submit it#

  1. Choose New → Run.
  2. Choose Edit as JSON (worker groups, drives, probes…).
  3. Replace the text in Run (JSON) with the contents of run.json.
  4. Choose Start run.

The console opens the run's page.

The New run form in JSON mode with the run specification

The form without JSON covers the same run: Name first-train, Image, Command python, GPUs 1, CPU cores per GPU 8, Memory per GPU 32Gi, and under More options a Time limit of 1h and the Environment STEPS=3000. It cannot add the script file: use JSON for that, or bake the script into your image.

astra astraeus run builds the same run from flags. --code . packs the current directory (with train.py) into the run and unpacks it in /workspace, where the command runs:

$ astra astraeus run --name first-train \
    --image nvcr.io/nvidia/pytorch:24.08-py3 \
    --gpus 1 --cpus 8 --mem 32G \
    --time 1h -e STEPS=3000 \
    --code . -- python train.py
first-train submitted to default: https://console.astralyx.cloud/o/my-org/w/default/jobs/default/first-train

With GPUs, --cpus and --mem are per GPU. astra astraeus run always sets restart_policy: Never. See the CLI reference for every flag.

Through the console's API, inside your workspace. <org>, <cluster>: your organisation's and cluster's short names (the workspace's Settings lists its clusters).

$ WS_API=https://console.astralyx.cloud/api/v1/orgs/<org>/workspaces/default/clusters/<cluster>/api
$ curl -sS -X POST "$WS_API/runs" \
    -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d @run.json | jq .metadata.name
"first-train"

The answer is 201 Created with the stored run (its metadata and spec). A name already taken is 409 JOB_ALREADY_EXISTS.

4. Follow its state#

A run goes Pending → Running → Completed. Its worker shows more detail: Pending (waiting for a machine), Preparing, Pulling (the image), Starting, Running, then Completed or Failed. See Run and worker states.

The run's page shows the state in its header, with the reason in words. Workers lists first-train-0 with its state, machine, GPUs and exit code. While it waits, Why it is not running yet says why, machine by machine.

The run's page with its worker running

$ astra astraeus status first-train
first-train: Running — All workers running
  first-train-0  Running  gpu-01
$ curl -sS "$WS_API/runs/first-train" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq .status.state
"Running"
$ curl -sS "$WS_API/workers?job=first-train" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq -r '.items[] | [.metadata.name, .status.state, .spec.assigned_node.name] | @tsv'
first-train-0   Running gpu-01

5. Read its log#

Logs stay on the machine. Each request asks the machine for the latest lines over the machine's own connection.

Open Logs on the run's page. Choose the worker in Log ·, tick Follow to keep it updating, Errors only to filter, Download to save it.

$ astra astraeus logs first-train-0 | tail -n 5
rank 0 of 1 on gpu-01: NVIDIA H100 80GB HBM3
step 200/3000 loss 2.3041 412 it/s
step 400/3000 loss 2.3027 431 it/s
step 600/3000 loss 2.3019 436 it/s
step 800/3000 loss 2.3020 437 it/s
$ curl -sS "$WS_API/workers/first-train-0/logs?tail=5" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq -r .logs

tail is the number of lines from the end: 1,000 by default, 10,000 at most.

The output above is illustrative: the rates depend on your GPU.

6. Stop it#

Deleting a run stops its workers and removes it. The machine stops each container, giving it 30 seconds after the stop signal before it is killed (the machine's stop grace), so a script can save a checkpoint.

On the run's page, choose Delete and confirm.

$ astra astraeus delete first-train
first-train deleted
$ curl -sS -X DELETE "$WS_API/runs/first-train" -H "Authorization: Bearer $ASTRAEUS_TOKEN" -o /dev/null -w '%{http_code}\n'
204

Warning

Deleting a run removes its workers and their logs. Download a log you want to keep first. Data written to a drive stays.

7. Run it again#

On the run's page, choose Run again. The form opens in JSON mode with the same specification under the name first-train-again. Change what you need and choose Start run.

To reuse a specification often, choose Save as template on the run's page, then pick it from the template list of New run.

Run the same astra astraeus run command again. A name can be reused once the earlier run is deleted.

POST the same run.json again, or change metadata.name to keep both.

Next steps#

  • Use more GPUs: set gpu_requests.count to 16 and let Astraeus pick the machines — see Multi-machine runs.
  • Read training data from a drive: see Drives.
  • Use a secret (an API key, a database password): see Credentials.