Recipes#
Each recipe builds one real thing from start to finish: the files, the commands, what you should see, how to check it worked and how to clean up. Start with the one closest to what you are doing and copy from it.
-
Train a PyTorch model on one GPU, follow its logs and keep the checkpoints on a drive.
Uses: runs, GPU requests, config files, drives kept on each machine, logs.
-
Distributed training across machines
torchrunon 2 machines × 8 GPUs over InfiniBand, with the rank and leader variables Astraeus sets and the NCCL settings it applies.Uses: gang start,
machines, topology rules, RDMA,RANK,WORLD_SIZE,MASTER_ADDR. -
A dataset on a drive, shared by runs
Download a dataset once per machine with a fill run, then mount it read-only in any number of runs. Runs go where the data already is.
Uses: drives kept on each machine, fills, data locations, read-only mounts, locality.
-
Preemptible fine-tuning with checkpoints
A low-priority fine-tune that saves a checkpoint on
SIGTERMand resumes where it stopped after a higher-priority run preempts it.Uses: priority, preemption, stop grace, a drive on one machine.
-
Hyperparameter sweep with an array
Twelve trials as one array run, each reading its index and writing its score to a shared drive, then a report run that ranks them.
Uses:
array,max_parallel,ASTRAEUS_ARRAY_TASK_ID, time limits and backfill. -
A schedule that starts a batch run every night at 01:30 in your time zone, and an alert when it fails.
Uses: schedules, time zones, time limits, alert rules.
-
Serve an internal service with an endpoint
A small HTTP API as a replica group behind an endpoint, called by name from another run; then external access and scale to zero.
Uses:
lifetime: Service, health checks, replica groups, endpoints, cluster DNS, ingress. -
Share a GPU fleet between teams
One cluster, two workspaces with their own quotas, pools, fair-share weights and maximum priorities, and what each team sees.
Uses: workspaces, cluster access terms, quotas, pools, fair share, preemption across workspaces.
-
Use cloud credentials without storing them
Read from S3 in a run with a credential that the machine fetches from AWS Secrets Manager with its own role. Nothing secret reaches the control plane.
Uses: credentials, machine identity,
secret_refs, outbound network policy.
How the recipes are written#
- Specifications are JSON. A run is the body of
POST /v1/runs(the older path/v1/jobsis the same collection); the console's New run → Edit as JSON takes the same body. Drives, schedules, endpoints, replica groups and credentials are JSON in the same way. Names are local to your workspace: writefood101, not the namespace. - Every step shows each way to do it when there is more than one, in
the order Console, CLI, API. The CLI is
astra: it covers runs (astra astraeus run,runs,logs,delete); drives, schedules, endpoints, replica groups and credentials are created in the console or with the API. -
The CLI works in a workspace you choose. Sign in once and pick the organisation, workspace and cluster:
$ astra login Open https://console.astralyx.cloud/device?code=KXQT-BRMW and approve code KXQT-BRMW Signed in as [email protected], working in acme/default $ astra use acme/vision@main Working in acme/vision -
API calls go through the console, with an API token, to your workspace's view of a cluster. Each recipe assumes these two variables:
$ export ASTRAEUS_TOKEN=$(cat ~/.astraeus-token) $ export API=https://console.astralyx.cloud/api/v1/orgs/acme/workspaces/vision/clusters/main/apiReplace
acme(organisation),vision(workspace) andmain(cluster) with yours. Create the token under Account → API tokens → New token and save it in~/.astraeus-token.
Related#
- Your first run
- Install the CLI
- Run specification
- Inside a worker: every variable a worker gets
- The queue
- Drives and data