Glossary#
The terms you meet in the console, astra and these pages, in alphabetical
order. Where the API, a specification field or a command uses another name,
it is given as API. The API accepts both the product names and the
original ones as paths (/runs and /jobs are the same collection).
| Term | API | Meaning |
|---|---|---|
| Activator | — | What, in the agent's edge part, holds a connection to a replica group scaled to zero while a replica starts, then hands it over. |
| Agent | — | The program on a machine that connects out to Astraeus: astraeus-agent, one binary whose parts run as services (astraeus-agent, astraeus-agent-drives, astraeus-agent-credentials, astraeus-agent-data, astraeus-agent-edge). Before October 2026 each part was its own binary (astraeus-worker, astraeus-dvagent, astraeus-secretsync, astraeus-catalog, astraeus-ingress). See How it works. |
| API token | ast_pat_… |
A personal token, created under Account & tokens, that acts as you through the console's API. |
| Array | spec.array |
A run whose template runs N times as independent workers, each with its index in ASTRAEUS_ARRAY_TASK_ID, at most max_parallel at once. See Arrays. |
| Astraeus Cloud | the hosted cluster | The shared cluster Astralyx operates for every organisation; each organisation is a tenant. |
| At risk | AtRisk |
A GPU's health when it shows signs that come before a fault. Avoided while healthy GPUs are free. |
| Backfill | — | Starting a waiting run early, in room held for the head of the queue, when its time limit makes it end before the head needs that room. See The queue. |
| Balanced endpoint | mode: balanced |
An endpoint with an Envoy listener in front of its backends. |
| Cache | evictable: true |
A drive kept on each machine whose unused copies may be removed when a disk fills. |
| Cluster | — | Where your work runs: the part of the Astralyx control plane that schedules it, and the machines enrolled in it. Astraeus Cloud (shared) or a cluster dedicated to an organisation; both operated by Astralyx. |
| Cluster access | binding | A workspace's access to one cluster, with its quota, weight, pools, max priority and host access. Creates the workspace's namespace there. |
| Console | — | The web interface and its API, https://console.astralyx.cloud. |
| Control plane | — | The Astralyx control plane (SaaS): the service, operated by Astralyx, that keeps your runs' desired state, places work on machines and serves the console. Machines connect out to it; it never connects to them. See How it works. |
| Copy | — | One machine's copy of a drive kept on each machine, in that machine's data location. |
| Cordon | nodes/<n>/cordon |
Stop new work from being placed on a machine; what runs there continues. |
| Credential | externalsecret (/external-secrets, /credentials) |
A reference to a secret in your own secret store, fetched on the machine with the machine's identity. See Credentials. |
| Data location | nodes/<n>/data-location |
The folder on a machine where it keeps copies of drives kept on each machine. Always chosen by a person. |
| Data source | connector (/connectors, /data-sources) |
A connection to data in object storage or on a filesystem, indexed on the machine that serves it. See Data sources. |
| Down | Down |
A machine whose heartbeat lapsed (60 s), or a worker on it. The work may still be running; it is re-adopted if the machine returns. |
| Drive | datavolume (/datavolumes, /drives) |
A name for data where it is: on one machine, on a shared filesystem, as scratch, or kept on each machine. See Drives and data. |
| Edge part | astraeus-agent-edge |
The part of the agent that programs Envoy for endpoints and external access on edge machines (before October 2026, the ingress agent, astraeus-ingress). |
| Endpoint | /endpoints |
A stable name for a set of workers chosen by labels, headless or balanced. See Endpoints. |
| Enrollment token | enrollment-tokens |
A single-use token in the install command; once used, it becomes the machine's identity. Valid 24 hours unused when made by the console. |
| Exposure | exposure |
Where a balanced endpoint's port opens: ingress, nodeport or both. |
| External access | external_accesses |
A port of a run published outside the cluster, through an endpoint named <run>-<access>. See External access. |
| Fabric | topology.astraeus.io/fabric |
An InfiniBand fabric: machines whose ports share a subnet. Found from the ports, or labelled. |
| Fair share | weight |
Ordering waiting runs of equal priority by how much of its quota each workspace holds, divided by its weight. |
| Gang | start: Gang |
A run whose workers are placed together and start together, or not at all. The default for runs that may span machines. |
| Head of the queue | — | The first waiting run that cannot start now; the queue holds room for it. |
| Headless endpoint | mode: headless |
An endpoint whose name resolves to every ready backend. |
| Heartbeat | — | A machine's or worker's sign of life, carried by its reports; when it lapses, the machine (60 s) or worker is marked Down. |
| Host access | host_access |
What a workspace's work may take from machines: host paths and privileged work. Nothing unless granted. |
| Idle | Idle |
A machine that registered and has not reported yet. |
| Lifetime | spec.lifetime |
Batch (ends when its workers end) or Service (runs until deleted). |
| Machine | node (/nodes, /machines) |
A computer running the Astraeus agent. See Machines. |
| Maintenance | Maintenance |
A machine taken out of service for maintenance. |
| Max priority | max_priority |
The highest run priority a workspace may ask for on a cluster (default 0). |
| Mesh | — | The WireGuard network (astraeus0, UDP 51820) that joins machines. |
| Mount | claim (/datavolume-claims, /mounts) |
An NFS mount of a drive's data from the machine that holds it to the machine of a worker that uses it. Made and removed automatically. |
| Namespace | namespace (ws-…) |
What a workspace is inside a cluster. Names inside it are qualified as <namespace>.<name>; you write local names. |
| NVLink domain | topology.astraeus.io/nvlink-domain |
Machines whose GPUs share NVLink (GB200 NVL72). Found from the GPUs. |
| On failure | spec.on_failure |
What a worker's failure restarts: RestartJob, RestartTask or nothing (FailJob). |
| Organisation | org | Who administers and pays: machines, workspaces, members, SSO, alerts. Roles owner, admin, member. |
| Placement | — | Astraeus's choice of machines for a run's workers: filters (fit, pools, health, network) and scores. See GPUs and placement. |
| Pool | a label, usually pool=<name> |
A group of machines granted to workspaces together. |
| Preemption | — | Stopping lower-priority work so a higher-priority run can start. Preempted workers are queued again without spending their restart budget. |
| Priority | spec.priority |
A run's place in the queue: higher first. Default 0. |
| Quota | quota |
The most a workspace may hold at once on a cluster: gpus, cpu_cores, memory_bytes, tasks (workers). |
| Rank | replica_index |
A worker's index in its run, from 0. In its environment as RANK. |
| Replica group | scalinggroup (/scaling-groups, /replica-groups) |
Identical service runs kept at a count, scaled on load and, behind an endpoint, to zero. See Replica groups. |
| Reservation | node-reservation (/node-reservations, /reservations) |
Machines set aside for a window, for some workspaces or for maintenance. See Reservations. |
| Restart policy | restart_policy |
Whether a failed worker is retried: OnFailure (default), Never, Always. |
| Run | job (/jobs, /runs) |
Work you ask for: an image, a command and what it needs. See Runs and workers. |
| Schedule | cronjob (/cronjobs, /schedules) |
Starts a run from a template on a cron schedule. See Schedules. |
| Shadow time | — | When the head of the queue will start at the latest, computed from running work's time limits. |
| Site | topology.astraeus.io/site |
Machines on one network. A run never spans sites. |
| Stale | Stale |
A running worker whose heartbeat is late; Down if it stays late. |
| Stop grace | --stop-grace (30 s) |
How long a machine waits after the stop signal before killing a worker's container. |
| Tenant | astraeus.io/tenant |
An organisation on a shared cluster. Its work runs only on its machines, its machines hear only its work. |
| Time limit | time_limit_seconds |
A worker is stopped after it and not restarted. Lets the queue backfill. |
| Weight | weight |
A workspace's fair-share weight on a cluster (default 1). |
| Worker | task (/tasks, /workers) |
One container of a run on one machine, named <run>-<rank>. |
| Worker group | task group (spec.task_groups) |
A named set of workers with its own template, in a run whose workers are not interchangeable. |
| Workspace | namespace, through a binding | A team's space: runs, drives, credentials, endpoints. One namespace on each cluster it has access to. See Organisations, workspaces and clusters. |