Skip to content

Glossary#

The terms you meet in the console, astra and these pages, in alphabetical order. Where the API, a specification field or a command uses another name, it is given as API. The API accepts both the product names and the original ones as paths (/runs and /jobs are the same collection).

Term API Meaning
Activator — What, in the agent's edge part, holds a connection to a replica group scaled to zero while a replica starts, then hands it over.
Agent — The program on a machine that connects out to Astraeus: astraeus-agent, one binary whose parts run as services (astraeus-agent, astraeus-agent-drives, astraeus-agent-credentials, astraeus-agent-data, astraeus-agent-edge). Before October 2026 each part was its own binary (astraeus-worker, astraeus-dvagent, astraeus-secretsync, astraeus-catalog, astraeus-ingress). See How it works.
API token ast_pat_… A personal token, created under Account & tokens, that acts as you through the console's API.
Array spec.array A run whose template runs N times as independent workers, each with its index in ASTRAEUS_ARRAY_TASK_ID, at most max_parallel at once. See Arrays.
Astraeus Cloud the hosted cluster The shared cluster Astralyx operates for every organisation; each organisation is a tenant.
At risk AtRisk A GPU's health when it shows signs that come before a fault. Avoided while healthy GPUs are free.
Backfill — Starting a waiting run early, in room held for the head of the queue, when its time limit makes it end before the head needs that room. See The queue.
Balanced endpoint mode: balanced An endpoint with an Envoy listener in front of its backends.
Cache evictable: true A drive kept on each machine whose unused copies may be removed when a disk fills.
Cluster — Where your work runs: the part of the Astralyx control plane that schedules it, and the machines enrolled in it. Astraeus Cloud (shared) or a cluster dedicated to an organisation; both operated by Astralyx.
Cluster access binding A workspace's access to one cluster, with its quota, weight, pools, max priority and host access. Creates the workspace's namespace there.
Console — The web interface and its API, https://console.astralyx.cloud.
Control plane — The Astralyx control plane (SaaS): the service, operated by Astralyx, that keeps your runs' desired state, places work on machines and serves the console. Machines connect out to it; it never connects to them. See How it works.
Copy — One machine's copy of a drive kept on each machine, in that machine's data location.
Cordon nodes/<n>/cordon Stop new work from being placed on a machine; what runs there continues.
Credential externalsecret (/external-secrets, /credentials) A reference to a secret in your own secret store, fetched on the machine with the machine's identity. See Credentials.
Data location nodes/<n>/data-location The folder on a machine where it keeps copies of drives kept on each machine. Always chosen by a person.
Data source connector (/connectors, /data-sources) A connection to data in object storage or on a filesystem, indexed on the machine that serves it. See Data sources.
Down Down A machine whose heartbeat lapsed (60 s), or a worker on it. The work may still be running; it is re-adopted if the machine returns.
Drive datavolume (/datavolumes, /drives) A name for data where it is: on one machine, on a shared filesystem, as scratch, or kept on each machine. See Drives and data.
Edge part astraeus-agent-edge The part of the agent that programs Envoy for endpoints and external access on edge machines (before October 2026, the ingress agent, astraeus-ingress).
Endpoint /endpoints A stable name for a set of workers chosen by labels, headless or balanced. See Endpoints.
Enrollment token enrollment-tokens A single-use token in the install command; once used, it becomes the machine's identity. Valid 24 hours unused when made by the console.
Exposure exposure Where a balanced endpoint's port opens: ingress, nodeport or both.
External access external_accesses A port of a run published outside the cluster, through an endpoint named <run>-<access>. See External access.
Fabric topology.astraeus.io/fabric An InfiniBand fabric: machines whose ports share a subnet. Found from the ports, or labelled.
Fair share weight Ordering waiting runs of equal priority by how much of its quota each workspace holds, divided by its weight.
Gang start: Gang A run whose workers are placed together and start together, or not at all. The default for runs that may span machines.
Head of the queue — The first waiting run that cannot start now; the queue holds room for it.
Headless endpoint mode: headless An endpoint whose name resolves to every ready backend.
Heartbeat — A machine's or worker's sign of life, carried by its reports; when it lapses, the machine (60 s) or worker is marked Down.
Host access host_access What a workspace's work may take from machines: host paths and privileged work. Nothing unless granted.
Idle Idle A machine that registered and has not reported yet.
Lifetime spec.lifetime Batch (ends when its workers end) or Service (runs until deleted).
Machine node (/nodes, /machines) A computer running the Astraeus agent. See Machines.
Maintenance Maintenance A machine taken out of service for maintenance.
Max priority max_priority The highest run priority a workspace may ask for on a cluster (default 0).
Mesh — The WireGuard network (astraeus0, UDP 51820) that joins machines.
Mount claim (/datavolume-claims, /mounts) An NFS mount of a drive's data from the machine that holds it to the machine of a worker that uses it. Made and removed automatically.
Namespace namespace (ws-…) What a workspace is inside a cluster. Names inside it are qualified as <namespace>.<name>; you write local names.
NVLink domain topology.astraeus.io/nvlink-domain Machines whose GPUs share NVLink (GB200 NVL72). Found from the GPUs.
On failure spec.on_failure What a worker's failure restarts: RestartJob, RestartTask or nothing (FailJob).
Organisation org Who administers and pays: machines, workspaces, members, SSO, alerts. Roles owner, admin, member.
Placement — Astraeus's choice of machines for a run's workers: filters (fit, pools, health, network) and scores. See GPUs and placement.
Pool a label, usually pool=<name> A group of machines granted to workspaces together.
Preemption — Stopping lower-priority work so a higher-priority run can start. Preempted workers are queued again without spending their restart budget.
Priority spec.priority A run's place in the queue: higher first. Default 0.
Quota quota The most a workspace may hold at once on a cluster: gpus, cpu_cores, memory_bytes, tasks (workers).
Rank replica_index A worker's index in its run, from 0. In its environment as RANK.
Replica group scalinggroup (/scaling-groups, /replica-groups) Identical service runs kept at a count, scaled on load and, behind an endpoint, to zero. See Replica groups.
Reservation node-reservation (/node-reservations, /reservations) Machines set aside for a window, for some workspaces or for maintenance. See Reservations.
Restart policy restart_policy Whether a failed worker is retried: OnFailure (default), Never, Always.
Run job (/jobs, /runs) Work you ask for: an image, a command and what it needs. See Runs and workers.
Schedule cronjob (/cronjobs, /schedules) Starts a run from a template on a cron schedule. See Schedules.
Shadow time — When the head of the queue will start at the latest, computed from running work's time limits.
Site topology.astraeus.io/site Machines on one network. A run never spans sites.
Stale Stale A running worker whose heartbeat is late; Down if it stays late.
Stop grace --stop-grace (30 s) How long a machine waits after the stop signal before killing a worker's container.
Tenant astraeus.io/tenant An organisation on a shared cluster. Its work runs only on its machines, its machines hear only its work.
Time limit time_limit_seconds A worker is stopped after it and not restarted. Lets the queue backfill.
Weight weight A workspace's fair-share weight on a cluster (default 1).
Worker task (/tasks, /workers) One container of a run on one machine, named <run>-<rank>.
Worker group task group (spec.task_groups) A named set of workers with its own template, in a run whose workers are not interchangeable.
Workspace namespace, through a binding A team's space: runs, drives, credentials, endpoints. One namespace on each cluster it has access to. See Organisations, workspaces and clusters.