Skip to content

CLI#

astra is Astralyx's command line. This page is its reference: every command, subcommand and flag, their defaults, the configuration file, environment variables and exit codes. To install it and sign in, see Install the CLI.

astra talks to the console at https://console.astralyx.cloud, as you, inside the organisation and workspace you choose. It signs in by approving the device in the browser, or uses a personal API token. Errors go to standard error with a non-zero exit status; see Exit codes.

astra#

Command Description
astra login Sign in: approve this device in the console
astra logout Forget the session
astra whoami Who you are, and where you work
astra orgs Your organisations and their workspaces
astra use Work in <org>/<workspace>, optionally on a cluster
astra astraeus … Runs, workers, logs, machines, GPUs, usage, templates
astra slurm … Slurm's commands, run on Astraeus
astra inference … Eos: models, deployments, API keys, chat

Global options: -h, --help prints help, on every command and subcommand; -V, --version prints the version.

Configuration#

astra keeps its settings in ~/.config/astra/config.json ($XDG_CONFIG_HOME/astra/config.json when XDG_CONFIG_HOME is set), created readable by you only (mode 0600) because it holds a session token.

~/.config/astra/config.json
{
  "url": "https://console.astralyx.cloud",
  "token": "ast_sess_…",
  "org": "acme",
  "workspace": "vision",
  "cluster": "lab-a"
}
Key Default Description
url https://console.astralyx.cloud The console; its API is <url>/api/v1.
token none A CLI session from astra login, or a personal API token.
org set by astra login (your first organisation) The organisation you work in.
workspace default after astra login The workspace you work in.
cluster the workspace's first cluster The cluster commands go to.

Environment variables override the file, when set and not empty:

Variable Overrides
ASTRA_URL url
ASTRA_TOKEN token. A personal API token works here, for CI.
ASTRA_ORG org
ASTRA_WORKSPACE workspace
ASTRA_CLUSTER cluster
ASTRA_IMAGE The container image for astra astraeus run --image, sbatch and srun

In CI, set ASTRA_TOKEN, ASTRA_ORG and ASTRA_WORKSPACE and skip astra login.

astra login#

Signs in by approving this device in the console (the OAuth device flow), and saves a CLI session valid for 30 days.

Flag Default Description
--url <URL> the configured url The console to sign in to; saved as url.
$ astra login
Open https://console.astralyx.cloud/device?code=KXQT-BRMW and approve code KXQT-BRMW
Signed in as [email protected], working in acme/default

Open the address, check that the code matches, and select Approve. The code expires after 10 minutes. If no organisation was chosen yet, login picks your first organisation and its default workspace.

astra logout#

Ends the session on the console and removes the token from the configuration file.

astra whoami#

Prints your e-mail, the console, and where you work.

$ astra whoami
[email protected] at https://console.astralyx.cloud
working in acme/vision on lab-a

astra orgs#

Lists your organisations and their workspaces.

$ astra orgs
ada  (Ada Lovelace (personal))  default
acme  (Acme Research)  default, nlp, vision

astra use#

$ astra use acme/vision@lab-a
Working in acme/vision
Argument Description
<org>/<workspace>[@<cluster>] The organisation and workspace to work in, and optionally the cluster. Without @<cluster>, commands use the workspace's first cluster. The workspace must exist and be yours.

astra astraeus#

Commands on the workspace's runs, workers and machines. They act through the console, as you, with your workspace role.

Command Alias Description
runs jobs The workspace's runs, on every cluster
status <run> One run: its state, why, and its workers
run [flags] -- <command> Start a run
workers <run> tasks A run's workers: rank, state, machine, why
logs <name> A run's or a worker's log
delete <run>… Delete runs; their workers stop
machines nodes The machines the workspace may use
gpus Their GPUs: model, use, health, who holds them
usage What the organisation's workspaces held, and the cost
templates The workspace's run templates

astra astraeus runs#

Flag Default Description
-a, --all off Also runs that ended (Completed, Failed, Cancelled) or were deleted.
$ astra astraeus runs
NAME                          CLUSTER       STATE       REASON
train                         lab-a         Running     All workers running
sweep                         lab-a         Pending     Waits for namespace quota: gpus 16 in use + 8 asked > 16

astra astraeus status#

$ astra astraeus status train
train: Running — All workers running
  train-0  Running  gpu-node-1
  train-1  Running  gpu-node-2

astra astraeus run#

Starts a run from flags, or from a run template.

Flag Default Description
--name <NAME> from the command: its file name in lower case, non-alphanumerics as -, then - and a number (python-48213) The run's name.
--image <IMAGE> $ASTRA_IMAGE The container image. Required unless --template is given.
-g, --gpus <N> 0 GPUs for the whole run; Astraeus picks the machines.
--machines <N> 0 (Astraeus decides) Exactly this many machines, the GPUs split evenly.
--cpus <N> 4 CPU cores: per GPU when the run has GPUs, otherwise per worker.
--mem <SIZE> 16G Memory, per GPU or per worker like --cpus. 512M, 64G, 16Gi, 1T; a plain number is MiB.
--time <DURATION> none A time limit: 90 (minutes), 2h, 1h30m, 45:30, 02:00:00, 1-12:00:00. At least 60 s.
-e, --env <KEY=value> none An environment variable. Repeatable.
--code <DIR> none Pack this directory's code with the run and unpack it in /workspace before the command runs. Needs a command. .git, target, node_modules, __pycache__, .venv, venv, .mypy_cache, .pytest_cache and .ipynb_checkpoints are left out; the pack must stay under 900 KiB. Put data on a drive.
--template <NAME> none Start from a run template of the workspace. The template's spec is used as it is: the resource flags above are ignored.
--pool <NAME> none Only machines labelled pool=<NAME>.
--model <MODEL> none A model of the workspace: its weights mounted read-only at /model, with ASTRAEUS_MODEL_PATH (and ASTRAEUS_MODEL_FILE for a single GGUF file) set.
--model-path <PATH> /model Where the model's weights appear. Requires --model.
--with-engine off Also serve the model beside the run: a worker group engine serves it, and the run's own group main gets OPENAI_BASE_URL and OPENAI_MODEL. The run ends when main does. Requires --model.
-w, --wait off Follow the run's log until it ends; exit non-zero unless it completed.
-- <command> [args…] the image's own The command and its arguments, after --.

Workers run with restart policy Never.

$ astra astraeus run --name train --image nvcr.io/nvidia/pytorch:24.08-py3 --gpus 16 --code . -- torchrun train.py
train submitted to lab-a: https://console.astralyx.cloud/o/acme/w/vision/jobs/lab-a/train

astra astraeus workers#

$ astra astraeus workers train
WORKER                        RANK   STATE       MACHINE               REASON
train-0                       0      Running     gpu-node-1
train-1                       1      Running     gpu-node-2

astra astraeus logs#

Argument or flag Default Description
<name> A worker (<run>-<rank>), or a run: its first worker's log.
-f, --follow off Keep printing new lines until the worker ends, polling every 2 s; then print its final state and exit non-zero unless it completed. While a run has no worker yet, -f waits up to 4 minutes for one.

astra astraeus delete#

astra astraeus delete <run> [<run>…] deletes each run (its workers stop) and prints <run> deleted.

astra astraeus machines#

$ astra astraeus machines
MACHINE               STATE     GPUS                CPU    POOL
gpu-node-1            Up        8× NVIDIA H100 80GB HBM3  96     h100
gpu-node-2            Up        8× NVIDIA H100 80GB HBM3  96     h100

astra astraeus gpus#

One line per GPU: <machine>/<index>, model, utilisation, health (Healthy, AtRisk, Faulty, Foreign), and free, held by <run>, or taken (held by another workspace).

astra astraeus usage#

Flag Default Description
--since <TIME> the first day of this month, 00:00 UTC Start of the period, RFC 3339. The period ends now.

Prints GPU-hours, core-hours and cost per workspace and cluster, and the total. See Usage, cost and budgets.

astra astraeus templates#

Lists the workspace's run templates: name and description.

astra slurm#

Slurm's commands, translated to Astraeus runs in the current workspace. astra slurm install-shims <DIR> makes sbatch, srun, squeue, scancel and sinfo symbolic links to astra in <DIR>; called by those names, astra behaves as that command. See Slurm compatibility.

Command Description
sbatch [options] <script> [args…] Submit a batch script. Its #SBATCH lines are read, then the options given before the script, which win. The script is added to the run as /astra/job.sbatch and run with bash, with the script's arguments. Prints Submitted batch job <name>.
srun [options] <command> [args…] Run a command as a run and follow its log until it ends.
squeue Runs waiting and running: JOBID (the run's name), ST (PD, CF, R), CLUSTER, REASON.
scancel <run>… Delete runs.
sinfo Machines by pool: PARTITION (the pool label), NODES, STATE (idle, drain when cordoned, down), NODELIST.
install-shims <DIR> Create the symbolic links in <DIR> (a directory on your PATH).

Options understood by sbatch and srun:

Option Becomes
-J, --job-name The run's name. Default: the script's file name (sbatch), or the command's name and a number (srun).
-N, --nodes N or MIN-MAX Exactly N machines, or between MIN and MAX.
-G, --gpus [type:]N N GPUs for the whole run.
--gpus-per-node, --gpus-per-task [type:]N N GPUs per machine.
--gres gpu[:type]:N N GPUs per machine. Other resources are not used.
-c, --cpus-per-task N Cores per machine; with GPUs, divided per GPU.
--mem SIZE Memory per machine (a plain number is MiB); with GPUs, divided per GPU.
--mem-per-cpu SIZE Memory per core, when --mem is not given.
-t, --time Time limit, in Slurm's forms (MM, MM:SS, HH:MM:SS, D-HH:MM:SS) or 2h.
-a, --array 0-99%10 An array of that size, at most %N at once. Also 1,2,3 (size 3) and N (size N + 1).
-p, --partition NAME Only machines with pool=NAME.
-w, --nodelist a,b Only these machines.
--container-image IMAGE The image (as with Pyxis). Without it, ASTRA_IMAGE; one of them is required.
--exclusive At least min(8, GPUs) GPUs on each machine.
--nice N Priority −N.

Accepted and ignored: -n/--ntasks, --ntasks-per-node (Astraeus runs one worker per machine), -o/--output, -e/--error (output goes to the run's log), -A/--account, -q/--qos (workspaces play that role), --export, -D/--chdir. Any other option is reported on standard error (astra: not used on Astraeus: …) and left out.

astra inference#

Commands for Eos, Astralyx's model serving; see Eos.

Command Flags
models --mine: only the workspace's models
library [words…] --category (repeatable: chat, code, reasoning, vision, embedding, tools, small), --publisher (repeatable), --size (repeatable: le3, 3-9, 9-20, 20-40, 40plus, or 7b), --engine (repeatable: llama.cpp, vllm), --license (permissive, restricted), --hide-gated, --machine (repeatable), --fits, --sort (popular, newest, smallest, name), --cluster, --json
show <family> --machine (repeatable), --cluster, --json
add <family:size-tag> --name, --credential, --cluster
pull <model> --machine (default: the one with most room)
deploy <model> --name, --cpu, --min (default 1; 0 scales to zero), --max (default 1), --context, --gpus (1, 2, 4 or 8)
deployments
keys create <name> --deployment (repeatable; none: all), --expires (RFC 3339)
keys list, keys revoke <name>
chat <deployment> --system, --temperature, --max-tokens

Exit codes#

Code Meaning
0 Success.
1 The command failed. The reason is on standard error, prefixed astra:. Also: --wait / --follow on a run or worker that ended other than Completed.
2 Invalid usage: an unknown command or flag, a missing argument, an invalid value.

Typical messages:

Message Meaning
astra: not signed in: \astra login` (or set ASTRA_TOKEN)` No token.
astra: the session is over or the token is wrong: \astra login`| The API answered401`.
astra: no organisation and workspace: \astra use /` (see `astra orgs`)| Runastra use`.
astra: the workspace has no cluster yet: add a machine in the console (it connects the workspace) The workspace has no cluster access.
astra: <message> (<CODE>) The API refused; see Errors.