Networking#
Astraeus gives every worker an address of its own, joins your machines with an encrypted WireGuard mesh, and serves names for workers, runs and endpoints inside the cluster — without any inbound connection from the control plane, and with only UDP 51820 open between machines. This page explains the model; the setup is in Network and firewalls and Names and service discovery.
Use these pieces when:
- Workers of a distributed run talk to each other: they reach the leader
at
MASTER_ADDR, and each other by name, across machines. - One run calls another: a training run calls an inference server, a client calls a database — by the other's name, inside the workspace.
- Several copies of a service need one stable name: an endpoint selects the ready workers and answers for them.
- A service must be reachable from outside the cluster: an external access on a run, or a balanced endpoint, opens a port on ingress machines or on the worker's machine.
- A service should cost nothing when idle: a replica group behind an endpoint scales to zero and wakes on the first request.
Addresses: a subnet per machine#
The cluster gives each machine a /24 subnet from the cluster's network
(10.240.0.0/16 by default). Each worker on a
machine gets its own address on that subnet, on a bridge the machine's agent
sets up, so two services on one machine never compete for a port.
The default network range holds 256 machines of up to 253 workers each. A dedicated cluster that needs more can use a larger range.
The mesh#
Machines are joined by WireGuard (interface astraeus0, UDP port 51820):
each machine has a peer for every other Up machine in its mesh, and a
route to that machine's subnet. Peers arrive over the connection the
machine opened, like everything else, so the mesh follows machines joining,
leaving and going down, with no call into the machine. Each machine makes its
WireGuard key itself and reports only the public half.
On a shared cluster such as Astraeus Cloud, each organisation has its own mesh: your machines never peer with another organisation's.
Host networking instead#
Some machines keep the host network: their workers use the machine's own addresses and ports instead of the mesh.
- A machine with RDMA (InfiniBand or RoCE) keeps the host network, which NCCL over RDMA needs.
- A machine where WireGuard cannot be set up falls back to it; the installer says so.
--network hostor--network meshat install chooses explicitly (meshfails rather than fall back).
Host-networked workers reach each other by machine address, and use the machine's resolver rather than the cluster's names.
Names#
Each machine's agent serves the cluster's names under the zone
astraeus.local, and points workers on the mesh at itself. It answers for:
| Name | Resolves to |
|---|---|
A worker, <run>-<rank> |
That worker's address |
A run, <run> |
Its workers' addresses |
An endpoint, <endpoint> |
Its ready backends |
Inside a worker, short names resolve within the workspace: train-0,
llm. The full form is <name>.<namespace>.astraeus.local, where the
namespace is the workspace's on that cluster (ws-…, shown in its
Settings). A machine resolves only its own organisation's names. Other
names go to the machine's own resolver. See Names and service
discovery.
Endpoints#
An endpoint gives a stable name to a set of workers, chosen by their labels within the workspace.
mode |
Behaviour |
|---|---|
headless (default) |
The name resolves to every ready backend; the client balances. Nothing in the path. |
balanced |
An Envoy listener on the ingress machines (and/or each backend's machine) in front of the ready backends, with health checks on health_check_path. |
A balanced endpoint's exposure decides where its port opens: ingress
(default: Envoy on machines running the agent's edge part), nodeport (on each
backend's machine, straight to the container) or both. Ports are allocated
from 30000–32767. An endpoint can name replica groups to wake when a
request finds no ready backend: the request is held while a replica starts.
See Endpoints and Replica groups and scale to
zero.
External access#
An external access is the short way to publish one port of a run: declare
it in the worker template, and Astraeus creates and maintains an endpoint
named <run>-<access> for the run's first worker, deleted with the run.
| Field | Description |
|---|---|
name |
Unique within the worker. |
target_port |
The container's port. |
external_port |
A port to request; empty to be allocated one (30000–32767). |
enabled |
Whether it is live. |
mode |
ingress, nodeport or both. |
port_env_name |
An environment variable the allocated port is put in. |
proxy |
An optional HTTP allow-list (methods, paths, body size, timeout) in front of the port. |
See External access.
Create one#
A Jupyter-style service on port 8888 reachable from outside, through the ingress machines:
- New → Run, Name
notebook, Imagequay.io/jupyter/base-notebook:latest, CPU cores2, Memory4Gi; under More options, Lifetime Service. -
Choose Edit as JSON (worker groups, drives, probes…) and add to
task_template: -
Start run. The endpoint
notebook-httpappears under Astraeus → Endpoints, with its port.
{
"metadata": {"name": "notebook"},
"spec": {
"lifetime": "Service",
"task_template": {
"image": "quay.io/jupyter/base-notebook:latest",
"requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296},
"external_accesses": [{"name": "http", "target_port": 8888, "enabled": true, "mode": "ingress"}]
}
}
}
An ingress access needs at least one machine running the agent's edge
part (installed with --agents drives,credentials,data,edge). The
astraeus CLI has no endpoint commands.