Skip to content

Machines#

A machine is a computer that runs the Astraeus agents and, on them, the workers of runs. It can be a GPU server in your rack, a cloud instance, a workstation, or a Mac serving models. The agents connect out to the cluster; nothing connects to the machine. In the API a machine is a node (/nodes, also /machines).

Add a machine when:

  • You have GPUs to use: a server with eight H100s, a cloud reservation, a workstation with an RTX or Radeon card.
  • You need CPU capacity: data preparation, evaluation, CPU-only simulations. A machine without a GPU runs CPU work.
  • You need an entry point for services: a machine running the agent's edge part publishes endpoints and holds requests while a replica group scaled to zero wakes.
  • You want to serve models on a Mac: a Mac with Apple silicon joins as an inference machine for Eos (see macOS).

What a machine is#

A Linux machine runs one agent, astraeus-agent, as a native program under systemd. Its main part (astraeus-agent.service) always runs; its drives, credentials and data parts (astraeus-agent-drives, astraeus-agent-credentials, astraeus-agent-data) by default; its edge part (astraeus-agent-edge) when asked. Only the workers of runs run in containers, on a containerd that ships with the agent (or the machine's Docker Engine). See How it works.

A machine takes its hostname as its name, unless the installer is given --name. Reinstalling it keeps its name and identity.

What it reports#

The worker reports what the machine has, and the scheduler places by it:

  • GPUs: model, memory, use, temperature and health, read from the NVIDIA driver (NVML) or AMD's amdgpu driver; their NVLink connections.
  • CPU and memory: cores, use, memory and pressure.
  • Disks and filesystems: local disks, RAID arrays and their health, shared filesystems — what the console suggests for drives and data locations.
  • Network: InfiniBand and RoCE ports, their speed and fabric.

The machine's page shows how its GPUs were read in the GPUs section: driver: healthy (the driver answered), nvidia-smi (read through the command-line tool), no-gpu, unresponsive or unavailable (the GPUs did not answer: the last known ones are kept).

States#

State Meaning New work
Idle Registered, has not reported yet. No
Up Reporting. Yes
Down Reports stopped: its heartbeat (60 s) lapsed. Its workers become Down too; they keep their GPUs until the machine returns and says whether they still run. No
Maintenance Taken out of service for maintenance; its reports do not bring it back. No

Cordoning a machine (Cordon on its page) stops new work from being placed there; what runs there continues. Uncordon reverses it. A machine under memory or disk pressure also takes no new work until it recovers.

See Run and worker states and Update, drain and remove.

Pools and topology#

Labels place machines. Two kinds matter:

  • Pools. A machine's pool is a label, usually pool=<name> (pool=h100), set when it is added. Workspaces are granted pools: their work runs only on machines carrying all the labels they were granted. See Organisations, workspaces and clusters.
  • Topology. Where a machine is, for placing runs that span machines:

    Label Meaning Set by
    topology.astraeus.io/site A site: a run never spans sites, because their machines talk over the internet The label, or the network the machine connects from
    topology.astraeus.io/rack A rack The label
    topology.astraeus.io/fabric An InfiniBand fabric Found from the subnet the ports report; the label wins
    topology.astraeus.io/nvlink-domain An NVLink domain (GB200 NVL72) Found from the GPUs

See Pools, labels and topology.

GPU health#

Every GPU has a health the scheduler acts on:

Health Meaning Placement
Healthy No sign of trouble. Used.
AtRisk Shows what comes before a fault: correctable memory errors growing, a link retrying, running hot. Avoided while healthy GPUs are free; never given to a run that asks for healthy GPUs only.
Faulty Reported a fault (an NVIDIA Xid, an AMD reset). Fenced: never used until an organisation admin clears the fault.
Foreign A process Astraeus did not start holds its memory. Not used meanwhile.

See GPUs.

Where a machine keeps data#

A machine's data location is the folder where it keeps copies of drives kept "on each machine". It is always a person's choice: the installer asks in a terminal (or takes --data-dir), or an organisation admin chooses it later on the machine's page. Until one is chosen nothing is written to the machine's disks, and a run that needs such a drive says Choose where to keep data on <machine>. See Drives and data.

Create one#

Adding a machine is two steps: get a single-use install command, then run it on the machine as root. You must be an organisation owner or admin.

  1. In a workspace, open Astraeus → Machines → Add machine.
  2. Set the Pool (pool=a100), or leave it empty.
  3. Choose Get the install command and copy it.
  4. Run it on the machine. The dialog shows Machine <name> connected. when it registers.

The Machines page of a workspace

Ask for an enrollment token; the answer carries the install command:

$ curl -sS -X POST https://console.astralyx.cloud/api/v1/orgs/acme/clusters/default/enrollment-tokens \
    -H "Authorization: Bearer $ASTRAEUS_TOKEN" -H 'content-type: application/json' \
    -d '{"labels": {"pool": "a100"}, "ttl_seconds": 86400}' | jq -r .install
echo '<token>' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --apiserver https://api.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme' --org-id <id> && rm -f ./astraeus-token

Run the printed command on the machine. labels are the machine's pool; ttl_seconds is how long the unused token stays valid.

On the machine, the installer prints what it does and starts the agent:

$ echo '<token>' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh \
    | sudo sh -s -- --apiserver https://api.astralyx.cloud --token-file ./astraeus-token --org 'Acme' --org-id <id>

Useful installer options: --name gpu-01, --data-dir /mnt/nvme0/astraeus, --agents drives,credentials,data,edge, --runtime docker, --network host. See Add a machine and the installer reference.