Skip to content

Network and firewalls#

Use this page to write firewall rules for your machines, to run them behind a proxy, or to understand how workers on different machines reach each other. The rule that shapes everything: a machine only connects out. Astraeus never opens a connection to a machine, so you never open a port on a machine for Astraeus to manage it.

flowchart LR
    subgraph site[Your network]
        A[Machine A] <-- "UDP 51820 (mesh)" --> B[Machine B]
    end
    A -- "HTTPS, outbound" --> CP["Astralyx control plane (SaaS)"]
    B -- "HTTPS, outbound" --> CP
    CP -. "work, on the same connection" .-> A
    A -- HTTPS --> R[Releases, registries, package repositories]

Outbound from every machine#

Destination Port When Used by
The Astraeus address (--apiserver in the install command shown in the Add machine dialog) TCP 443 Always The agent and all its parts: registration, the connection that brings the machine its work (a long-lived HTTPS response), reports every 10 seconds, logs and commands.
The console (https://<console>/api/v1/install.sh) TCP 443 When you run the install command curl, to fetch the installer
The releases address: https://console.astralyx.cloud/releases (--releases in the console's command), otherwise github.com and the download hosts it redirects to (*.githubusercontent.com) TCP 443 Install, and Update from the console The installer; the machine agent for updates
The distribution's package repositories 80/443 When the installer installs iptables, iproute2, wireguard-tools, SELinux tools or a GPU driver apt, dnf, yum, zypper, pacman
Driver repositories: developer.download.nvidia.com and dl.fedoraproject.org (RHEL family), mirrors.rpmfusion.org (Fedora), deb.debian.org and security.debian.org (Debian), repo.radeon.com (AMD, Ubuntu) 80/443 Only when the installer installs a driver from them The installer
Container registries your runs use (docker.io, ghcr.io, nvcr.io, your own) 443 When a worker's image is pulled astraeus-containerd (or Docker)
ghcr.io 443 Agent runs in an OpenShell sandbox (pulls the pinned supervisor image) astraeus-containerd
Secret managers your credentials reference (AWS, Google Cloud, Azure, Oracle Cloud, HashiCorp Vault, …) and their identity endpoints 443 When a worker on the machine uses the credential astraeus-agent-credentials
Object stores your data sources point at 443 When a data source is served from this machine astraeus-agent-data
Whatever your runs fetch (model hubs, datasets, APIs) Any At run time, subject to the run's outbound policy The workers

Machines need no access to the console after installation, except for Update, which downloads from the releases address.

Between machines#

Port Protocol When Why
51820 UDP Machines on the mesh WireGuard: workers on different machines reach each other by address.
2049 TCP A drive one machine serves to workers on another NFS 4.2, from the agent's drives part (astraeus-agent-drives).
20049 RDMA The same, when both ends share an RDMA network NFS over RDMA. If it fails, the client falls back to TCP and says so.
Any TCP/UDP — Multi-machine runs on the host network (machines with RDMA, or --network host) NCCL and the run's own rendezvous (MASTER_PORT is 29500) talk directly between machines.

On the mesh, worker-to-worker traffic between machines is carried inside WireGuard: only UDP 51820 needs to be open between them.

Inbound, only for services you expose#

Nothing listens for Astraeus. Ports open on a machine only when you expose a service:

Port On When
The endpoint's node port The machine running the backend worker An endpoint with exposure: nodeport or both. Published on every interface, or on ASTRAEUS_NODE_PORT_ADDRESS when set.
The endpoint's listener ports Machines running the ingress agent and Envoy exposure: ingress or both.
8800 (plain HTTP) Machines running the ingress agent The inference gateway, unless turned off (ASTRAEUS_GATEWAY_ADDRESS empty). Put your own TLS in front.

See Endpoints and External access.

What never listens on the machine's network:

  • containerd: Unix socket /run/astraeus/containerd/containerd.sock only, root only; no TCP, debug or metrics listener.
  • The machine's DNS server (:53) and tool gateway (:8801): bound to the node network's gateway address on the workers' bridge, not to the machine's own interfaces.

Network modes#

A Linux machine runs its workers in one of two modes, chosen at install:

Mode What workers get Chosen when
Mesh (ASTRAEUS_NETWORK_MODE=bridge, ASTRAEUS_MESH=true) Each worker has its own address on the machine's subnet; workers on other machines are reached over WireGuard; names resolve through the machine's DNS. Two deployments never share a port. The default when WireGuard works. Force it with --network mesh (the installer fails rather than falls back).
Host (ASTRAEUS_NETWORK_MODE=host) Workers share the machine's network stack and its ports. The machine has RDMA (InfiniBand or RoCE: NCCL over RDMA needs the host network), WireGuard cannot be set up, or --network host.

The installer prints its choice:

· network: each task gets an address of its own, the machines are joined by WireGuard (UDP 51820)
· this machine has RDMA: tasks use the host network (what NCCL over RDMA needs); --network mesh to override

A Mac always uses the host network. See macOS.

Node IPAM#

Astraeus gives each machine its own subnet for workers, so no two machines ever share one:

Setting Default
Cluster network 10.240.0.0/16
Subnet per machine /24

With the defaults the cluster has 256 machine subnets of 254 worker addresses each. A machine's subnet is released when it is removed. The gateway (.1) of each subnet is the machine's bridge, where its DNS server listens. If the cluster network overlaps your own networks, tell Astralyx before you add machines to a dedicated cluster.

On the machine:

Interface What
astraeus0 The WireGuard interface.
astraeus-cni0 The bridge of the machine's subnet (the bundled containerd).
astraeus-cni1 The bridge of the default worker network (the bundled containerd).
astraeus-br The bridge of the Docker network astraeus (with --runtime docker).
astraeus-* network namespaces in /run/netns One per worker (the bundled containerd).

The agent turns on IPv4 forwarding and adds iptables rules: a MASQUERADE for each worker subnet towards the outside (the mesh's own traffic is exempted), FORWARD accept rules for its bridges, inserted ahead of any DROP (for example Docker's), and DNAT rules for published ports.

The WireGuard mesh#

  • Each machine generates its WireGuard key once (/var/lib/astraeus/worker/mesh/private.key, mode 0600) and reports only the public key.
  • Its peers are every other machine of the cluster that is not Down (on a shared cluster, only its own organisation's). Each peer is its public key, its address on port 51820, and its subnet as the only allowed source. Keepalives every 25 seconds keep NAT mappings open.
  • Peers arrive over the machine's own outbound connection, so the mesh follows machines joining, leaving and going down without Astraeus ever connecting to a machine.

A machine's address is the source address of its route to the Astraeus address, unless you set it. Machines must reach each other at those addresses on UDP 51820. Machines on different private networks (two offices, home and a datacentre) cannot reach each other's private addresses: give each a reachable address with ASTRAEUS_NODE_IP in /etc/astraeus/agent.env, then sudo systemctl restart astraeus-agent.

/etc/astraeus/agent.env (excerpt)
ASTRAEUS_NODE_IP=203.0.113.24

Note

The installer rewrites /etc/astraeus/agent.env on every run. Add ASTRAEUS_NODE_IP again after you rerun it, or set it in a systemd drop-in (see Behind an HTTP proxy for the pattern).

Multi-machine runs never span sites, so traffic between sites over the mesh is for services, not for collectives.

DNS on the machine#

On the mesh, each machine serves the cluster's DNS zone, astraeus.local, on its subnet's gateway (port 53):

  • It answers A queries for running workers (<worker>), runs (<run>) and endpoints (<endpoint>), within the worker's workspace first.
  • It answers NXDOMAIN for other names in the zone, and forwards everything else to the machine's own resolver.
  • Workers on the bridge get it as their resolver, with <workspace>.astraeus.local and astraeus.local as search domains.
  • Workers on the host network keep the machine's resolver and reach each other by machine address.

It also enforces a run's outbound policy: a name the policy does not allow is refused. See Names and service discovery.

firewalld#

When firewalld is running (Fedora, RHEL and its rebuilds, openSUSE), the installer configures it, permanently and live:

What Why
Zone astraeus, target ACCEPT, with interfaces astraeus0, astraeus-cni0, astraeus-cni1, astraeus-br (moved from any zone they were in) Workers reach the machine's DNS and tool gateway, which firewalld sees as incoming traffic. Without it, runs fail with Could not resolve host.
Port 51820/udp in the default zone The mesh.
Policy astraeus-forwarding (firewalld 0.9 or newer): ingress zone ANY, egress zone astraeus, target ACCEPT Forwarded traffic towards workers: published ports and the mesh.

Interfaces that do not exist yet join the zone when they appear. --uninstall removes the zone, the policy and the port.

$ sudo firewall-cmd --get-active-zones
astraeus
  interfaces: astraeus0 astraeus-cni0
public
  interfaces: eth0

Other host firewalls#

The installer configures firewalld only. With another firewall that drops incoming traffic by default (ufw, a hand-written nftables ruleset), allow what firewalld's setup allows: traffic coming in on the astraeus bridges and the mesh interface, and UDP 51820. With ufw:

$ sudo ufw allow in on astraeus-cni0
$ sudo ufw allow in on astraeus-cni1
$ sudo ufw allow in on astraeus0
$ sudo ufw allow 51820/udp

A run's outbound policy is enforced inside each worker's own network namespace (an nftables table inet astraeus, or iptables when nft is missing) and does not touch the machine's rules.

Behind an HTTP proxy#

There is no proxy option in the installer. Give each piece the standard variables:

  1. The installer. The first curl runs as you; the installer's own downloads run under sudo, which drops your environment. Pass the proxy through sudo env:

    $ export https_proxy=http://proxy.example.com:3128
    $ curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo env https_proxy=$https_proxy sh -s -- --apiserver https://api.astralyx.cloud --token-file ./astraeus-token
    

    Package managers use their own proxy settings (Acquire::http::Proxy for apt, proxy= in dnf.conf).

  2. The agent and containerd. The agent's services honour HTTPS_PROXY, HTTP_PROXY and NO_PROXY for their connection to Astraeus; it is plain HTTPS and works through proxies. containerd uses the same variables to pull images. Set them in systemd drop-ins, which survive reruns of the installer:

    set-proxy.sh
    for u in containerd agent agent-drives agent-credentials agent-data; do
      sudo mkdir -p /etc/systemd/system/astraeus-$u.service.d
      sudo tee /etc/systemd/system/astraeus-$u.service.d/proxy.conf >/dev/null <<'EOF'
    [Service]
    Environment=HTTPS_PROXY=http://proxy.example.com:3128
    Environment=HTTP_PROXY=http://proxy.example.com:3128
    Environment=NO_PROXY=localhost,127.0.0.1,10.240.0.0/16
    EOF
    done
    sudo systemctl daemon-reload
    sudo systemctl restart astraeus-containerd astraeus-agent astraeus-agent-drives astraeus-agent-credentials astraeus-agent-data
    
  3. Runs. Workers do not inherit the agent's environment. A run that needs the proxy sets the variables itself.

Update from the console does not use the proxy

The machine agent downloads a release for Update without a proxy. Behind a proxy that is the only way out, update by running the install command again instead (Update, drain and remove).

The agent works through a proxy that tunnels (CONNECT). Behind a proxy that terminates TLS, the machine authenticates with its token instead of its certificate.