Skip to content

Networking#

Astraeus gives every worker an address of its own, joins your machines with an encrypted WireGuard mesh, and serves names for workers, runs and endpoints inside the cluster — without any inbound connection from the control plane, and with only UDP 51820 open between machines. This page explains the model; the setup is in Network and firewalls and Names and service discovery.

Use these pieces when:

  • Workers of a distributed run talk to each other: they reach the leader at MASTER_ADDR, and each other by name, across machines.
  • One run calls another: a training run calls an inference server, a client calls a database — by the other's name, inside the workspace.
  • Several copies of a service need one stable name: an endpoint selects the ready workers and answers for them.
  • A service must be reachable from outside the cluster: an external access on a run, or a balanced endpoint, opens a port on ingress machines or on the worker's machine.
  • A service should cost nothing when idle: a replica group behind an endpoint scales to zero and wakes on the first request.

Addresses: a subnet per machine#

The cluster gives each machine a /24 subnet from the cluster's network (10.240.0.0/16 by default). Each worker on a machine gets its own address on that subnet, on a bridge the machine's agent sets up, so two services on one machine never compete for a port.

The default network range holds 256 machines of up to 253 workers each. A dedicated cluster that needs more can use a larger range.

The mesh#

Machines are joined by WireGuard (interface astraeus0, UDP port 51820): each machine has a peer for every other Up machine in its mesh, and a route to that machine's subnet. Peers arrive over the connection the machine opened, like everything else, so the mesh follows machines joining, leaving and going down, with no call into the machine. Each machine makes its WireGuard key itself and reports only the public half.

On a shared cluster such as Astraeus Cloud, each organisation has its own mesh: your machines never peer with another organisation's.

Host networking instead#

Some machines keep the host network: their workers use the machine's own addresses and ports instead of the mesh.

  • A machine with RDMA (InfiniBand or RoCE) keeps the host network, which NCCL over RDMA needs.
  • A machine where WireGuard cannot be set up falls back to it; the installer says so.
  • --network host or --network mesh at install chooses explicitly (mesh fails rather than fall back).

Host-networked workers reach each other by machine address, and use the machine's resolver rather than the cluster's names.

Names#

Each machine's agent serves the cluster's names under the zone astraeus.local, and points workers on the mesh at itself. It answers for:

Name Resolves to
A worker, <run>-<rank> That worker's address
A run, <run> Its workers' addresses
An endpoint, <endpoint> Its ready backends

Inside a worker, short names resolve within the workspace: train-0, llm. The full form is <name>.<namespace>.astraeus.local, where the namespace is the workspace's on that cluster (ws-…, shown in its Settings). A machine resolves only its own organisation's names. Other names go to the machine's own resolver. See Names and service discovery.

Endpoints#

An endpoint gives a stable name to a set of workers, chosen by their labels within the workspace.

mode Behaviour
headless (default) The name resolves to every ready backend; the client balances. Nothing in the path.
balanced An Envoy listener on the ingress machines (and/or each backend's machine) in front of the ready backends, with health checks on health_check_path.

A balanced endpoint's exposure decides where its port opens: ingress (default: Envoy on machines running the agent's edge part), nodeport (on each backend's machine, straight to the container) or both. Ports are allocated from 30000–32767. An endpoint can name replica groups to wake when a request finds no ready backend: the request is held while a replica starts. See Endpoints and Replica groups and scale to zero.

External access#

An external access is the short way to publish one port of a run: declare it in the worker template, and Astraeus creates and maintains an endpoint named <run>-<access> for the run's first worker, deleted with the run.

Field Description
name Unique within the worker.
target_port The container's port.
external_port A port to request; empty to be allocated one (30000–32767).
enabled Whether it is live.
mode ingress, nodeport or both.
port_env_name An environment variable the allocated port is put in.
proxy An optional HTTP allow-list (methods, paths, body size, timeout) in front of the port.

See External access.

Create one#

A Jupyter-style service on port 8888 reachable from outside, through the ingress machines:

  1. New → Run, Name notebook, Image quay.io/jupyter/base-notebook:latest, CPU cores 2, Memory 4Gi; under More options, Lifetime Service.
  2. Choose Edit as JSON (worker groups, drives, probes…) and add to task_template:

    "external_accesses": [{"name": "http", "target_port": 8888, "enabled": true, "mode": "ingress"}]
    
  3. Start run. The endpoint notebook-http appears under Astraeus → Endpoints, with its port.

notebook.json
{
  "metadata": {"name": "notebook"},
  "spec": {
    "lifetime": "Service",
    "task_template": {
      "image": "quay.io/jupyter/base-notebook:latest",
      "requested_resources": {"cpu_cores": 2, "memory_bytes": 4294967296},
      "external_accesses": [{"name": "http", "target_port": 8888, "enabled": true, "mode": "ingress"}]
    }
  }
}
$ curl -sS -X POST "$WS_API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @notebook.json
$ curl -sS "$WS_API/runs/notebook/external-accesses" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '.items[] | {name, external_host, external_port, status}'

An ingress access needs at least one machine running the agent's edge part (installed with --agents drives,credentials,data,edge). The astraeus CLI has no endpoint commands.