Skip to content

How machines reach each other#

Workers on different machines talk to each other over the mesh: WireGuard between every two machines of the cluster, on UDP 51820 (Network and firewalls). In a data centre every machine reaches every other at its address. At home it is not that simple: a house sits behind its router's NAT, a carrier may add its own (CGNAT), and a Windows PC runs its machine inside WSL, behind the PC's own NAT. This page says how machines find a direct path anyway, how to see which machines reach which, and what Astraeus does when two cannot.

Direct, or not at all

Traffic between two machines always goes directly between them, encrypted end to end by WireGuard. It never goes through Astralyx: Astralyx only passes addresses between the machines. There is no relay today, so two machines that cannot reach each other directly do not talk at all — and Astraeus keeps work that talks together on machines that do.

Direct, not reached#

For every other machine, each machine says one of:

State Means
direct A WireGuard handshake in the last 3 minutes: the two machines reach each other directly. Shown with the round trip over the tunnel (direct · 3 ms), measured every 5 minutes.
not reached Every address the other machine gave was tried, without an answer. Shown with why, as far as Astraeus can tell (different networks, a NAT that changes its port, a firewall).
trying Its addresses are being tried, one at a time.
unknown Nothing said yet: one of the two is down, runs an agent older than this, or does not join the mesh (a machine on the host network, a Mac).

A handshake is mutual: if either machine sees one, both reach each other. If either has tried everything without one, they do not.

How a machine is found#

Each machine tells Astralyx where it may be reached, best first, and refreshes it every 10 minutes (every 2 while a machine is being tried):

Address What it is Who tries it first
Its local network The machine's own IPv4 addresses, with the network each is on (192.168.68.122 on 192.168.68.0/22). Machines on the same network.
Its PC's address (WSL) Under WSL, the Windows PC's own addresses on its networks, read from Windows. WSL's packets leave from there. Machines on the PC's network.
IPv6 The machine's global IPv6 addresses — stable ones only, not the temporary (privacy) ones that change every day. Machines that have IPv6 too.
Its public endpoint Where the internet sees its UDP port 51820 (its router's public address and port), and the kind of NAT it is behind. Machines in other networks.

The address a machine was registered with (its machine address) is always among them, so a cluster in one network works as it always has.

Another machine tries these addresses one at a time — WireGuard keeps one address per machine — for 30 seconds each, then the next, and starts over after the last; after three rounds without an answer it moves on every 2 minutes. As soon as a handshake comes, the address is kept where the other machine's packets came from. That matters behind a NAT: a WSL machine says an address inside its PC, nobody else can reach that address, and its packets arrive from the PC's address and a port the PC chose. The machine it reached answers there, and keeps answering there.

The public endpoint#

To learn its public endpoint, the machine asks two public STUN servers (stun.l.google.com:19302 and stun.cloudflare.com:3478) from its own UDP port 51820. Each answers with the address and port it saw. Only that question and its answer go to them, nothing else.

The two answers The machine says
One of the machine's own addresses No NAT: its address is public.
The same public address and port Behind a NAT that keeps its port whatever the destination: machines in other networks can reach it by hole punching.
Two different ports Behind a symmetric NAT, which gives each destination a new port: no machine in another network can reach it directly.

To use other STUN servers, or none, set ASTRAEUS_STUN_SERVERS in /etc/astraeus/agent.env (comma-separated host:port; empty: not asked), then sudo systemctl restart astraeus-agent. Without a public endpoint, a machine is still reached on its local network and over IPv6.

Machines on the same network#

Two machines in one house reach each other on their local network, whatever NAT is between the house and the internet.

A Windows PC under WSL (WSL's default networking, NAT) is the case to know: its machine is inside the PC, behind the PC's own NAT, and nothing from the house reaches it first. It works because the PC's machine reaches the others: it sends first, through the PC's NAT, the other machine answers where the packets came from, and keepalives every 25 seconds keep the PC's NAT open. Both machines need a current agent — with an older one their Reaches says unknown; Update it from the machine's page: an older agent set the other machine's address back every minute, cutting the PC off again.

WSL's mirrored networking#

With mirrored networking, WSL shares the PC's network: the machine has the PC's own address on the house's network, and the other machines reach it directly, like any Linux machine. It needs Windows 11 22H2 or later and the current WSL. The installer offers it on a Windows PC, as one of its Windows steps:

  1. networkingMode=mirrored under [wsl2] in your %USERPROFILE%\.wslconfig (a backup is kept).
  2. A Hyper-V firewall rule, Astralyx mesh, that lets UDP 51820 in to WSL (needs an administrator; if Windows refuses, the installer prints the PowerShell command to run as administrator).

It takes effect after wsl --shutdown (or a reboot). The installer does not turn it on with --yes alone when Docker Desktop is installed, since Docker Desktop's own networking may need NAT; answer the question to choose. To do it yourself:

%USERPROFILE%\.wslconfig
[wsl2]
networkingMode=mirrored
New-NetFirewallHyperVRule -Name 'Astralyx mesh' -DisplayName 'Astralyx mesh' -Direction Inbound -VMCreatorId '{40E0AC32-46A5-438A-A0B2-2B479E8F2E90}' -Protocol UDP -LocalPorts 51820 -Action Allow
wsl --shutdown

Undo it by removing the line from .wslconfig (or restoring the .wslconfig.astraeus-backup file next to it) and running Remove-NetFirewallHyperVRule -Name 'Astralyx mesh'.

Machines in different networks#

Two homes, a home and an office: neither reaches the other's local address. In this order:

  1. IPv6. Many home connections have IPv6, and IPv6 has no NAT — only the router's firewall, which lets the answer to a machine's own packets in. When both machines have a global IPv6 address, they try it first. If your router has IPv6 turned off, turning it on is often all it takes.
  2. Hole punching. When two machines have tried everything without an answer, and both are behind NATs that keep their port, Astralyx tells both — at the same moment — to send to the other's public endpoint for a minute. Each router lets the other machine's packets in because its own machine just sent to that address and port, and the tunnel comes up directly. A pair is punched again at most every 10 minutes while it is not reached.

What hole punching cannot get through:

  • A symmetric NAT on either side (some carriers' CGNAT, some business routers): the port the other machine would need is not the one anyone can learn. The machine says so (Behind a symmetric NAT).
  • Two machines behind one public address but on different local networks (two networks of one office): their local addresses are the way; check the routing between those networks.
  • A firewall that drops UDP to or from the internet.

Then, for that pair, either:

  • Forward UDP 51820 on one machine's router to that machine. It is then reached at its public endpoint by everyone.
  • Or set ASTRAEUS_NODE_IP to an address the others can reach (Network and firewalls).

Astraeus does not open ports on your router (UPnP, NAT-PMP or PCP) and has no relay: both would be choices for your organisation to make, and neither is offered yet.

See who reaches whom#

A machine's page has a Reaches section: each other machine, direct · 3 ms or not reached (with why), since when, and — folded — where the machine may be reached and the NAT it is behind.

Machines that cannot reach each other are a finding on the Topology page: fedora and majin cannot reach vitor: different networks; work that talks together is kept on machines that reach each other, with a fix (Topology → Findings).

$ astra astraeus machines show majin
majin  Up
GPUs: 1× NVIDIA GeForce RTX 4060
Address: 172.19.156.94  Windows (WSL2)  agent 0.1.0

Reaches
  fedora               direct · 3 ms  since 2026-10-09 10:02
  vitor                not reached — different networks, and their NATs or firewalls let no direct connection through  since 2026-10-09 10:04
Reached at (behind a NAT that keeps its port (machines elsewhere can punch through)):
  172.19.156.94:51820          local network (172.19.144.0/20)
  192.168.69.50:51820          its PC's address (WSL) (192.168.68.0/22)
  189.1.2.3:40211              public endpoint

--json prints the machine and its row as the API returns them.

Method and path What
GET /machines/{name}/reach One machine: reaches (each other machine: state — direct, not_reached, trying, unknown — since, rtt_ms, why, the addresses in use), candidates and nat.
GET /topology/reach Every pair of the workspace's machines (pairs), the groups that reach each other (groups), and each machine's candidates and nat.

Both need topology:read, and show only the machines in the workspace's pools.

get_machine_topology includes reaches and where the machine is reached; get_topology includes the finding (Connect your AI assistant).

Work that talks to other work#

A worker that calls a service of the workspace over the network — a function calling a search service, an indexer writing to a vector store — must run on a machine that reaches that service's machine. Two rules say so:

Rule Places the worker When to use it
near On the same machine as the service. The service is small and the two belong together: one machine always reaches itself.
reaches On any machine that reaches the service's machine (or on that machine). near would be too strict: the service needs a GPU elsewhere, or there is no room next to it.
requested_resources:
  cpu_cores: 1
  memory_bytes: 268435456
  reaches: [embed]      # an Eos deployment of the workspace

A machine that is not reached from where the service runs is left out, and the run says why when no machine is left: Calls embed over the network, and no machine it may use reaches fedora, where it runs (vitor does not). A pair that is unknown (an older agent, no mesh) counts as reached.

In a template, a function or run whose part needs a service part (a replica group, a deployment, a long-running run, or an endpoint in front of one) gets reaches for it, unless it already asks for near.

Troubleshooting#

What you see What to do
A Windows PC and a Linux machine in one house are not reached Update both agents (Update on each machine's page). If it stays so, turn on mirrored networking, and check that no firewall on the Linux machine drops UDP 51820 (Network and firewalls).
Two machines on one network are not reached, neither under WSL A firewall drops UDP 51820 on one of them, or the network isolates its clients (a guest Wi-Fi).
Machines in two homes are not reached, both behind a NAT that keeps its port Hole punching is tried every 10 minutes; a router that drops unasked UDP after only a few seconds can defeat it. Forward UDP 51820 on one router, or turn on IPv6.
A machine is behind a symmetric NAT Forward UDP 51820 to it on its router, give both machines IPv6, or ask the carrier for a public IPv4 address.
unknown for every pair The agents are older than this: update them.
No public endpoint shown The STUN servers did not answer (UDP to the internet blocked), or ASTRAEUS_STUN_SERVERS is empty.