How it works#
This page explains what runs where, how your machines and Astraeus talk, and what happens between submitting a run and its completion. Read it once before you operate machines or debug a run: most behaviour follows from one rule — machines dial out; nothing dials in.
The pieces#
flowchart LR
U["You<br/>Console · CLI · API"]
CP["Astralyx control plane (SaaS)"]
subgraph M["Your machine"]
W["astraeus-agent"]
A["its parts: drives · credentials<br/>data · edge"]
end
U -- "HTTPS, as you,<br/>inside your workspace" --> CP
W -- "HTTPS, opened by the machine" --> CP
A -- "HTTPS, opened by the machine" --> CP
There are three parts:
- You — the console at
https://console.astralyx.cloud, theastraCLI and the API. Every request is made as you, inside a workspace, and is checked against your role in that workspace. - The Astralyx control plane (SaaS) — operated by Astralyx. It keeps the desired state of your runs, places workers on machines, runs the queue, tracks the health of machines and workers, and serves the console. Your cluster is either Astraeus Cloud (shared) or a cluster dedicated to your organisation; both are provided and operated by Astralyx.
- Your machines — the computers that run your workloads, with the Astraeus agent installed.
The agent on a machine#
A machine runs one native program, astraeus-agent, under systemd: each of
its parts is a service of its own. Only the workers of runs run in
containers, on a containerd that ships with the agent (or the machine's
Docker Engine with --runtime docker).
| Part (service) | Installed | What it does |
|---|---|---|
Main (astraeus-agent) |
Always | Runs workers in containers, gives them GPUs, runs health probes, collects metrics and logs, sets up the machine's subnet, the WireGuard mesh and the cluster DNS. |
Drives (astraeus-agent-drives) |
By default | Exports and mounts drives over NFS when a worker runs on another machine than its data. |
Credentials (astraeus-agent-credentials) |
By default | Fetches credentials from your secret stores with the machine's own identity and writes them for the worker, readable by root only. |
Data (astraeus-agent-data) |
By default | Indexes the data behind data sources (files, Parquet footers, profiles) and keeps the index on the machine. |
Edge (astraeus-agent-edge) |
With --agents …,edge |
Programs Envoy for endpoints and external access, and holds connections while a replica group scaled to zero wakes. |
See Add a machine and the installer reference.
Machines dial out#
Every part of the agent connects out to the control plane over HTTPS, on connections the machine opens. Over them it receives the work that concerns this machine and this part only — nothing about other machines' work — and reports back what happened: state changes of its workers, heartbeats, the machine's status and metrics. Reports are sent on a regular interval, and at once when a worker changes state.
When Astraeus needs something from a machine on demand — a worker's log, a file listing of a data source, an HTTP request to a worker — it hands the request to the machine over the same outbound connection, and the machine answers. Nothing ever opens a connection to the machine.
| Part | Receives | Reports |
|---|---|---|
| Main | Workers placed on this machine (with their drives, world size and the address of rank 0), its mesh peers, the cluster DNS names, on-demand requests | Worker states and health, heartbeats, machine info, metrics, answers |
| Drives | Drive mounts this machine is an end of | Mount progress |
| Credentials | The credential references this machine's workers need | Sync results — never values |
| Data | Data sources assigned to this machine, on-demand requests | Data source health, index summaries |
| Edge | Endpoints and replica groups, and their ready backends | Traffic, wake requests |
This has practical consequences:
- One firewall rule. A machine needs outbound HTTPS to the Astraeus address shown in the console (and to the console itself, for the installer and releases). Machines in the WireGuard mesh also reach each other on UDP 51820. No inbound port is ever opened for Astraeus. See Network and firewalls.
- Corporate proxies work. The connections are ordinary HTTPS requests and streams that survive HTTP proxies.
- A machine that loses its network keeps running. Its containers continue; when it reconnects, it re-adopts them and reports their state.
Identity#
A machine joins with a single-use enrollment token from the console. When it joins, it receives its own credential and authenticates with it from then on. A machine can act only for itself: it receives only its own work and can report only on its own workers. You can revoke a machine at any time from the console. See Security model.
Liveness#
Astraeus separates "is this machine alive" from "what is its status".
| What | Lifetime | Renewed by | When it lapses |
|---|---|---|---|
| Machine heartbeat | 60 s | Every report; a keep-alive every 20 s | The machine becomes Down, and its workers become Down with it. |
| Worker heartbeat | The worker's heartbeat TTL (default 60 s) | Every report | A running worker gets 15 s of grace, then becomes Stale, then Down after a further TTL. |
Down and Stale mean Astraeus lost sight of a worker, not that it
stopped. The worker keeps its GPUs, cores, memory and its share of the
workspace's quota until the machine comes back and says whether the
container still runs (it is re-adopted) or not (it fails with Container not
found, and everything is released). After an interruption on the Astralyx
side, missing heartbeats are not counted as losses for one heartbeat period,
so control-plane maintenance does not mark your machines Down.
See Run and worker states for every state.
The life of a run#
sequenceDiagram
autonumber
actor You
participant CP as Astralyx control plane (SaaS)
participant M as Machine
You->>CP: Submit the run (console, astra or API)
CP->>CP: Check your role, validate, check priority
CP-->>You: 201 Created — run Pending
CP->>CP: Create workers, queue, quota, placement
CP-->>M: Worker placed on this machine
M->>M: Credentials, drives, image pull
M->>CP: Preparing, Pulling, ReadyToStart
Note over M: A gang waits until every worker is ready
M->>CP: Starting, Running (heartbeats)
You->>CP: Logs
CP-->>M: Request the log
M->>CP: Log lines
CP-->>You: Log lines
M->>CP: Completed (exit code 0)
CP->>CP: Run Completed
- Submission. You create a run from the console,
astraor the API, inside a workspace. Astraeus checks your role, that every reference stays inside the workspace, that the priority is within what the workspace is granted, and that the specification is valid. The run isPending. - Workers. Astraeus turns the run into workers named
<run>-<rank>(<run>-<group>-<rank>with worker groups). When you asked for GPUs without a machine count, it also picks the shape: how many machines, and how many GPUs on each. - The queue. Waiting runs are ordered by priority, then fair share, then submission time. A run over its workspace's quota waits without blocking others. See The queue.
- Placement. Astraeus filters the machines the workspace may use and scores the rest. Placing a worker reserves its GPUs, cores and memory and re-checks the quota at the same moment, so two decisions can never overcommit a machine or a quota.
- On the machine. The machine receives the worker over its outbound
connection. It fetches credentials (through the agent's credentials part),
mounts drives (through its drives part), pulls the image and creates the container. A gang's
workers wait at a barrier (
ReadyToStart) until all of them are ready, then start together. - Running. The machine reports state changes as they happen, and heartbeats and metrics on every report.
- Completion. When a container exits, the machine reports
Completed(exit code 0) orFailed. The run's failure policy decides what happens next — restart the worker, restart the whole run, or fail. When every worker of a batch run has completed, the run isCompleted.
What is kept where#
| Where | What |
|---|---|
| Kept by Astralyx | Your account, sessions and API tokens (hashed), organisations, members, workspaces and their cluster access, run templates, usage and the audit log. For each cluster: runs, workers, machines and their latest status, drives (definitions, not data), credential references, endpoints, replica groups, schedules, roles, the history of every state change and the events derived from it, and metric series of machines, GPUs, workers and drives. Machine credentials are kept as hashes. |
| Your machines | Containers and their logs, the data in drives and their copies, credential values (written readable by root only, removed when no worker on the machine needs them), data source indexes and profiles, the machine's WireGuard private key. |
Logs stay on the machine: the console fetches them from the machine when you open them. See What leaves your machines for the full account of what crosses the boundary.