Skip to content

Submit a run#

A run is a container image, a command and what it needs (GPUs, CPU, memory, data), started on your workspace's machines. This page shows how to start one in each of the three ways — console, CLI and API — and how to set everything a typical run needs: the image and its registry, the command and its environment, resources, a time limit and what happens on failure. It ends with what happens after you submit, and how to stop, delete or repeat a run.

For the complete list of fields, see the run specification. For several machines, see Multi-machine runs.

Before you begin#

  • You are an editor or admin of a workspace that has access to at least one cluster with a machine. Viewers can read runs but not start them. See Organisations, workspaces and clusters.
  • For the CLI: install astra, sign in and choose the workspace (Install the CLI):

    $ astra login
    Open https://console.astralyx.cloud/device?code=QWRT-PLKM and approve code QWRT-PLKM
    Signed in as [email protected], working in acme/default
    $ astra use acme/vision
    Working in acme/vision
    
  • For the API: an API token from the console (Account → API tokens → New token). The examples below put it in ASTRA_TOKEN.

Start a training run#

The example trains a ResNet-50 on one GPU, reading ImageNet from the drive imagenet, with a 12-hour limit. The console and API examples run code kept on the drive; the CLI example sends the code in the current directory with the run.

  1. Open the workspace, then Runs → New run (/o/<org>/w/<workspace>/new-run).
  2. Choose the Cluster, and give the run a Name: resnet50. Lower-case letters, digits and -.
  3. Image: nvcr.io/nvidia/pytorch:24.08-py3.
  4. Command: python. Arguments: /data/code/train.py --epochs 90 (separated by spaces; the code is on the drive mounted in step 6).
  5. GPUs: 1. With GPUs, the next two fields are CPU cores per GPU (8) and Memory per GPU (64Gi).
  6. Drives: imagenet:/data mounts the drive imagenet at /data.
  7. Open More options and set Time limit to 12h.
  8. Select Start run. The console opens the run's page; its state is Pending until a machine takes it.

The New run form with the fields above filled in

The panel on the right shows what the workspace has on this cluster: its Quota, Weight, Max priority and Pools.

--code . packs the current directory (code only, at most 900 KiB) and unpacks it in /workspace, where the command runs.

$ astra astraeus run --name resnet50 \
    --image nvcr.io/nvidia/pytorch:24.08-py3 \
    --gpus 1 --cpus 8 --mem 64G --time 12h \
    --code . -- python train.py --epochs 90
resnet50 submitted to gpu-east: https://console.astralyx.cloud/o/acme/w/vision/jobs/gpu-east/resnet50
$ astra astraeus runs
NAME                          CLUSTER       STATE       REASON
resnet50                      gpu-east      Running     1 of 1 workers running

Add --wait to follow the log until the run ends; the command then exits non-zero if the run did not complete.

astra astraeus run has no flag for drives or credentials. For a run that mounts the imagenet drive, save the spec as a template and start it with --template, or use the API.

The workspace's view of a cluster is https://console.astralyx.cloud/api/v1/orgs/<org>/workspaces/<workspace>/clusters/<cluster>/api, the cluster's /v1 with your workspace's names.

resnet50.json
{
  "metadata": { "name": "resnet50" },
  "spec": {
    "task_template": {
      "image": "nvcr.io/nvidia/pytorch:24.08-py3",
      "command": "python",
      "args": ["/data/code/train.py", "--epochs", "90"],
      "time_limit_seconds": 43200,
      "datavolume_refs": [{ "name": "imagenet", "mount_path": "/data" }],
      "requested_resources": {
        "gpu_requests": { "count": 1 },
        "per_gpu": { "cpu_cores": 8, "memory_bytes": 68719476736 }
      }
    }
  }
}
$ API=https://console.astralyx.cloud/api/v1/orgs/acme/workspaces/vision/clusters/gpu-east/api
$ curl -sS -X POST "$API/jobs" \
    -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H 'Content-Type: application/json' \
    --data-binary @resnet50.json | jq -r .metadata.name
resnet50

The answer is 201 Created with the run as created. /runs is accepted as an alias of /jobs.

From a spec file#

Keep runs you repeat as files next to your code. The API takes JSON; YAML works if you convert it as you send it.

  1. In New run, select Edit as JSON (worker groups, drives, probes…). The form's run appears as the JSON body.
  2. Paste or edit your spec. Names are local to the workspace: write imagenet, not a namespace.
  3. Select Start run.

astra has no command that submits a file. Use the API (next tab), or keep the spec as a workspace template and start it with --template (see Run again and templates).

$ yq -o=json run.yaml | curl -sS -X POST "$API/jobs" \
    -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H 'Content-Type: application/json' --data-binary @-

A field the cluster does not know is refused, never ignored: 400 UNKNOWN_FIELD with fields this API does not know: spec.replicas.

Name the run#

  • Letters, digits, - and _, no - at either end, at most 57 characters (less with worker groups). The console accepts lower-case only; use lower-case everywhere.
  • Unique in the workspace on that cluster: a second run with the same name is refused with 409 JOB_ALREADY_EXISTS. Delete the old one, or pick another name.
  • Workers are named after the run: resnet50-0, resnet50-1… (<run>-<group>-<rank> with worker groups).
  • astra astraeus run without --name names the run after the command: python-48213.

Choose the workspace and cluster#

A run lives in one workspace, on one cluster. The workspace's quota, pools and maximum priority on that cluster apply to it.

The Cluster field of New run lists the clusters the workspace has access to. The runs list (Runs) shows every cluster's runs together.

astra use <org>/<workspace> chooses the workspace; astra use <org>/<workspace>@<cluster> also fixes the cluster. Without a cluster, astra uses the workspace's first. ASTRA_ORG, ASTRA_WORKSPACE and ASTRA_CLUSTER override the saved choice for one command.

The organisation, workspace and cluster are in the path: /api/v1/orgs/<org>/workspaces/<workspace>/clusters/<cluster>/api/jobs.

Images and private registries#

image is any OCI image reference; busybox means docker.io/library/busybox:latest. pull_policy decides when the machine pulls it:

pull_policy Behaviour
IfNotPresent (default) Pull only if the image is not on the machine. Use a new tag for a new build.
Always Pull every time the worker starts.
None Never pull; the worker fails with Image <image> is not present and the pull policy is None if it is missing.

To pull from a private registry:

  1. Create a credential whose value is JSON with the registry login:

    {"username": "robot$ci", "password": "<token>", "server": "registry.example.com"}
    
  2. Reference it from the run:

    "task_template": {
      "image": "registry.example.com/nlp/train:2026-09",
      "registry_secret_ref": { "external_secret_name": "registry-login" }
    }
    

The machine fetches the credential itself; Astralyx never sees its value. Until the credential is on the machine, the worker stays in Preparing with waiting for registry secret registry-login to be synced to this node. A wrong name or login ends the worker with ErrImagePull: <error>.

Command, arguments and environment#

  • command is one executable. It replaces the image's ENTRYPOINT and CMD; args are its arguments. Without command, the image's ENTRYPOINT runs with args (or the image's CMD).
  • There is no shell. "args": ["--nodes=$WORLD_SIZE"] passes the literal text $WORLD_SIZE. To expand variables, run a shell:

    "command": "bash",
    "args": ["-c", "python train.py --rank $RANK --world-size $WORLD_SIZE"]
    
  • env sets variables. Yours override the ones Astraeus injects (RANK, MASTER_ADDR…); see Inside a worker.

  • Secrets do not go in env, which anyone who can read the run can see. Reference a credential instead; the machine resolves it:

    "secret_refs": [
      { "external_secret_name": "hf-token", "env_vars": { "token": "HF_TOKEN" } },
      { "external_secret_name": "gcs-key", "mount_path": "/var/secrets/gcs" }
    ]
    

In the console, Environment takes KEY=value pairs (comma-separated or one per line) and Credentials takes credential names. With astra, use -e KEY=value (repeatable).

Resources#

Every worker states its CPU and memory, in one of three forms:

Form Fields When to use it
Per worker cpu_cores, memory_bytes CPU runs, arrays, and GPU runs on a fixed number of machines.
Per GPU per_gpu.cpu_cores, per_gpu.memory_bytes GPU runs: a worker with 8 GPUs gets 8 times the amount, whatever shape the scheduler picks. The console and astra use this form whenever you ask for GPUs.
Share of a machine node_fraction (0–1] A fixed share of whichever machine it lands on. The container then has no CPU or memory limit.
  • Memory is a hard limit. A worker that exceeds it is killed: OOMKilled: the container exceeded its memory limit.
  • Cores are a hard quota unless you set cpu_burst: true, which lets a worker use idle cores while keeping its share under contention.
  • gpu_requests.count is the run's total; the scheduler splits it across machines. Choose GPU models, memory and health in GPUs and placement.
  • Units: the console reads 512Mi, 64Gi and 64G as binary sizes (64 GiB) and a plain number as bytes; astra reads 64G and 64Gi as 64 GiB and a plain number as MiB; the API takes bytes (68719476736 is 64 GiB).

Time limits#

time_limit_seconds (60 s to 366 days) bounds how long each worker runs, counted from when it is Running. Past it, the worker is stopped — SIGTERM, then SIGKILL after its machine's stop grace (30 s by default) — and the run fails with Worker resnet50-0: Exceeded its time limit of 12h. A timed-out worker is never restarted.

Set one whenever you know roughly how long a run takes: a run with a time limit can backfill into room held for a larger run, so it often starts sooner.

More options → Time limit: 90m, 2h, 1h30m, 1d, or a number of seconds.

--time 12h. astra also reads Slurm's forms: 30 (minutes), 02:00:00, 1-12:00:00.

"time_limit_seconds": 43200 in task_template.

Failures and restarts#

Three settings decide what a failure does. The defaults suit a single-worker batch run: a worker that exits non-zero is restarted (after 10 s, doubling up to 5 min), at most 10 times.

Setting Values Default
task_template.restart_policy OnFailure, Always, Never OnFailure
spec.on_failure RestartJob (all workers), RestartTask (the failed one), FailJob (no restart; the run fails) RestartJob for a gang, else RestartTask
spec.lifetime Batch (runs to completion), Service (restarted whenever it ends, until deleted) Batch

The CLIs do not restart

astra astraeus run and astra slurm sbatch set restart_policy: Never: a failed worker fails its run at once. The console and the API keep the default, OnFailure.

For a run that should never repeat work (it writes results as it goes), use on_failure: FailJob. For a server or a notebook, use lifetime: Service. See Run and worker states for the exact back-off and budget.

What happens after you submit#

  1. The run is accepted as Pending (Run created) and Astraeus creates its workers.
  2. The queue orders waiting runs by priority, then fair share, then submission time, and places workers on machines that pass every filter. A run that waits says why: see Read a pending reason.
  3. Each placed worker's machine prepares it (Preparing), pulls the image (Pulling), starts it (Starting) and reports Running.
  4. The run ends Completed when every worker exits 0, Failed when a worker fails for good, or Cancelled when its workers were stopped.

A run's page: state, workers, machines and GPUs

Follow it on the run's page, with astra astraeus status resnet50, or with GET $API/jobs/resnet50. See Logs, metrics and debugging.

Run again and templates#

  • Run again on a run's page opens New run with that run's specification as JSON, named <run>-again. Change what you need and select Start run.
  • Save as template (on a run's page, or in New run) keeps the specification in the workspace under a name; saving under an existing name replaces it.
  • In New run, Start from a template… lists the workspace's templates; Use it loads one as JSON, Remove deletes it.
$ astra astraeus templates
llama-70b finetune  16 H100, 24 h, checkpoints on /ckpt
$ astra astraeus run --template "llama-70b finetune" --name finetune-oct
finetune-oct submitted to gpu-east: https://console.astralyx.cloud/o/acme/w/vision/jobs/gpu-east/finetune-oct

With --template, the template's specification is used as it is; --image, --gpus and the other resource flags are ignored.

Templates belong to the workspace, not the cluster:

$ curl -sS "https://console.astralyx.cloud/api/v1/orgs/acme/workspaces/vision/templates" \
    -H "Authorization: Bearer $ASTRA_TOKEN" | jq -r '.items[].name'
llama-70b finetune
$ curl -sS -X POST "https://console.astralyx.cloud/api/v1/orgs/acme/workspaces/vision/templates" \
    -H "Authorization: Bearer $ASTRA_TOKEN" -H 'Content-Type: application/json' \
    -d '{"name": "resnet50", "description": "1 GPU, 12 h", "spec": '"$(jq .spec resnet50.json)"'}'

A template's spec is checked by the cluster only when a run is made from it. Templates are at most 256 KiB. DELETE …/templates/<id> removes one.

To start a run on a timetable, use a schedule.

Stop or delete a run#

Deleting a run stops its workers (each gets SIGTERM, then SIGKILL after the stop grace) and removes the run, its workers and their logs.

Deleting removes the logs

Download what you need from the run's Logs tab first. There is no undo.

On the run's page select Delete, type the run's name to confirm, and select the confirm button. The dialog says how many workers will be stopped and how many GPUs freed.

$ astra astraeus delete resnet50
resnet50 deleted

Several names delete several runs.

$ curl -sS -X DELETE "$API/jobs/resnet50" -H "Authorization: Bearer $ASTRA_TOKEN" -w '%{http_code}\n'
204

To stop a run but keep its record and logs, stop its workers instead: POST $API/tasks/<worker>/stop (optional body {"reason": "…"}). A stopped worker ends Cancelled and is not restarted; once every worker has ended, the run is Cancelled (All workers ended; some were cancelled). There is no console button or CLI command for this.

astra astraeus run flags#

Flag Default Description
--name from the command The run's name.
--image $ASTRA_IMAGE The container image. Required unless --template.
-g, --gpus 0 GPUs for the whole run; the scheduler picks the machines.
--machines 0 Exactly this many machines, GPUs split evenly.
--cpus 4 Cores — per GPU when there are GPUs, else per worker.
--mem 16G Memory — per GPU when there are GPUs, else per worker. 64G, 512M; a plain number is MiB.
--time none Time limit: 2h, 1h30m, 90 (minutes), 02:00:00, 1-00:00:00.
-e, --env — KEY=value, repeatable.
--code — Pack this directory (at most 900 KiB packed; .git, target, node_modules, __pycache__, .venv, venv, .mypy_cache, .pytest_cache, .ipynb_checkpoints skipped) and run the command in /workspace. Needs a command.
--template — Start from a workspace template.
--pool — Only machines labelled pool=<value>.
--model — Mount a model of the workspace at /model.
--model-path /model Where the model is mounted. Needs --model.
--with-engine off Serve the model beside the run. Needs --model.
-w, --wait off Follow the first worker's log until the run ends.
-- <command> [args…] the image's The command and its arguments.

astra sets restart_policy: Never and does not set start: a run on several machines starts its workers independently. For a gang, use the console, the API, or a template whose spec sets start: Gang.

Troubleshooting#

Symptom Cause Fix
400 UNKNOWN_FIELD: fields this API does not know: spec.replicas A field the API does not have. Workers come from machines, node_selection.count or array; see the run specification.
400 VALIDATION_ERROR with detail A value out of range or a contradiction. Each detail entry names the field and the rule.
400 NAME_NOT_DNS_SAFE The name (or a derived worker name) is not a DNS label. Shorter name, letters, digits and - only.
403 PRIORITY_NOT_ALLOWED: priority 50 is above namespace …'s maximum of 0 The workspace may not ask for that priority. Lower it, or ask an admin to raise the maximum.
409 JOB_ALREADY_EXISTS The name is taken. Delete the old run or pick another name.
astra: the workspace has no cluster yet: add a machine in the console The workspace has no cluster. Add a machine.
The run stays Pending No room, quota, or rules that no machine meets. Read the pending reason.