Skip to content

Recipes#

Each recipe builds one real thing from start to finish: the files, the commands, what you should see, how to check it worked and how to clean up. Start with the one closest to what you are doing and copy from it.

  • Train on one GPU machine


    Train a PyTorch model on one GPU, follow its logs and keep the checkpoints on a drive.

    Uses: runs, GPU requests, config files, drives kept on each machine, logs.

  • Distributed training across machines


    torchrun on 2 machines × 8 GPUs over InfiniBand, with the rank and leader variables Astraeus sets and the NCCL settings it applies.

    Uses: gang start, machines, topology rules, RDMA, RANK, WORLD_SIZE, MASTER_ADDR.

  • A dataset on a drive, shared by runs


    Download a dataset once per machine with a fill run, then mount it read-only in any number of runs. Runs go where the data already is.

    Uses: drives kept on each machine, fills, data locations, read-only mounts, locality.

  • Preemptible fine-tuning with checkpoints


    A low-priority fine-tune that saves a checkpoint on SIGTERM and resumes where it stopped after a higher-priority run preempts it.

    Uses: priority, preemption, stop grace, a drive on one machine.

  • Hyperparameter sweep with an array


    Twelve trials as one array run, each reading its index and writing its score to a shared drive, then a report run that ranks them.

    Uses: array, max_parallel, ASTRAEUS_ARRAY_TASK_ID, time limits and backfill.

  • Nightly batch job


    A schedule that starts a batch run every night at 01:30 in your time zone, and an alert when it fails.

    Uses: schedules, time zones, time limits, alert rules.

  • Serve an internal service with an endpoint


    A small HTTP API as a replica group behind an endpoint, called by name from another run; then external access and scale to zero.

    Uses: lifetime: Service, health checks, replica groups, endpoints, cluster DNS, ingress.

  • Share a GPU fleet between teams


    One cluster, two workspaces with their own quotas, pools, fair-share weights and maximum priorities, and what each team sees.

    Uses: workspaces, cluster access terms, quotas, pools, fair share, preemption across workspaces.

  • Use cloud credentials without storing them


    Read from S3 in a run with a credential that the machine fetches from AWS Secrets Manager with its own role. Nothing secret reaches the control plane.

    Uses: credentials, machine identity, secret_refs, outbound network policy.

How the recipes are written#

  • Specifications are JSON. A run is the body of POST /v1/runs (the older path /v1/jobs is the same collection); the console's New run → Edit as JSON takes the same body. Drives, schedules, endpoints, replica groups and credentials are JSON in the same way. Names are local to your workspace: write food101, not the namespace.
  • Every step shows each way to do it when there is more than one, in the order Console, CLI, API. The CLI is astra: it covers runs (astra astraeus run, runs, logs, delete); drives, schedules, endpoints, replica groups and credentials are created in the console or with the API.
  • The CLI works in a workspace you choose. Sign in once and pick the organisation, workspace and cluster:

    $ astra login
    Open https://console.astralyx.cloud/device?code=KXQT-BRMW and approve code KXQT-BRMW
    Signed in as [email protected], working in acme/default
    $ astra use acme/vision@main
    Working in acme/vision
    
  • API calls go through the console, with an API token, to your workspace's view of a cluster. Each recipe assumes these two variables:

    $ export ASTRAEUS_TOKEN=$(cat ~/.astraeus-token)
    $ export API=https://console.astralyx.cloud/api/v1/orgs/acme/workspaces/vision/clusters/main/api
    

    Replace acme (organisation), vision (workspace) and main (cluster) with yours. Create the token under Account → API tokens → New token and save it in ~/.astraeus-token.