Names and service discovery#
Every running worker, every run and every endpoint has a DNS name under astraeus.local. Workers on the node network resolve these names with the cluster DNS that each machine serves, so a run finds its peers, its leader and the services it uses by name, without knowing which machine they landed on.
Use these names to:
- point the workers of a distributed run at their leader (
train) or at a given rank (train-3); - call a service from another run (
redis,llm) through its endpoint; - configure a client once, while the workers behind the name move between machines.
The names#
Inside a workspace, every name is <name>.<namespace>.astraeus.local. The namespace is the workspace's own on the cluster: ws- and 12 hexadecimal digits, shown in the workspace's settings.
| Name | Resolves to | While |
|---|---|---|
<worker>.<namespace>.astraeus.local |
The worker's address | The worker is Running |
<run>.<namespace>.astraeus.local |
The run's leader's address. If the leader is not running: every running worker of the run | At least one worker of the run is Running |
<endpoint>.<namespace>.astraeus.local |
The addresses of the endpoint's ready workers (several answers) | At least one is ready |
There are no other names: no per-group names and no -headless names.
Worker names follow the run's name:
| Run | Workers |
|---|---|
| Without worker groups | <run>-0, <run>-1, … (the index is the worker's rank) |
| With worker groups | <run>-<group>-0, <run>-<group>-1, … |
The leader is worker 0 of the leader group (spec.leader_group, by default the first group). Members of a replica group are runs named <group>-<index>. A replica group's load balancer is an endpoint named like the group, so <group> resolves to its ready members.
Endpoints that a run's external_accesses create (<run>-<access>) are not published in DNS: they point at a worker that has a name already.
How resolution works#
flowchart LR
W["Worker (node network)"] -->|"queries 10.240.3.1:53"| D["Cluster DNS on the worker's machine"]
D -->|"name under astraeus.local"| Z["Answers from the zone (TTL 5 s)"]
D -->|"any other name"| R["The machine's own resolver"]
- Each machine serves the cluster DNS on its node network's gateway, the first address of the machine's
/24(for example10.240.3.1, port 53). Each machine receives the zone from Astraeus, over the connection it keeps open, whenever the zone changes. It answers from memory: a lookup never leaves the machine. - Names under
astraeus.localare answered withArecords (IPv4) with a TTL of 5 seconds, so a client sees a moved worker quickly. A name in the zone that does not exist, or has no address right now, getsNXDOMAIN. Names are case-insensitive. - Every other name is forwarded to the machine's own resolver. A run with an outbound policy (
network.egresswithdefault: deny) resolves only the names its policy allows. See Networking. - A machine knows the names of every workspace of its own organisation, and never the names of another organisation's workspaces.
Each worker on the node network gets its own /etc/resolv.conf:
nameserver 10.240.3.1
search ws-3f9c2a7d41be.astraeus.local astraeus.local
options ndots:2
Because of the search list, short names work within the workspace: train-0 means train-0.ws-3f9c2a7d41be.astraeus.local. A name of another workspace of your organisation needs its namespace: api.ws-8d01e44b7c2a (reaching it is also subject to the run's outbound policy).
A resolv.conf you put in the worker yourself (a config file mounted at /etc/resolv.conf) takes precedence.
Examples#
Resolve a peer from inside a worker:
Point every worker of a distributed run at the leader by name. The run is named train (16 GPUs on 2 machines, one worker on each), so train resolves to its leader:
{
"metadata": {"name": "train"},
"spec": {
"task_template": {
"image": "ghcr.io/acme/trainer:5.1",
"command": "sh",
"args": ["-c", "torchrun --nnodes=$WORLD_SIZE --node_rank=$RANK --nproc_per_node=8 --master_addr=train --master_port=29500 train.py"],
"requested_resources": {"cpu_cores": 32, "memory_bytes": 274877906944, "gpu_requests": {"count": 16}, "machines": 2}
}
}
}
Every worker also gets the leader's address directly in ASTRAEUS_LEADER_ADDRESS and MASTER_ADDR, which works on the host network as well. See Inside a worker.
Call a service from another run through its endpoint:
Workers on the host network#
Workers on the host network do not use the cluster DNS: they keep the machine's resolver, and astraeus.local names do not resolve there. Machines with RDMA (InfiniBand or RoCE) run their workers on the host network by default, because NCCL over RDMA needs it, and so do machines where WireGuard cannot be set up.
On the host network:
- use
ASTRAEUS_LEADER_ADDRESS(orMASTER_ADDR) to reach the leader; - workers reach each other at their machines' addresses;
- to reach a service, use its endpoint's address from the endpoint's
runtime.backends, or publish it with anodeportexposure.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
NXDOMAIN for a worker or run |
It is not Running yet, or the name is misspelt. |
Check the worker's state. Names follow the run: <run>-<index> or <run>-<group>-<index>. |
NXDOMAIN for an endpoint |
No worker is ready (not running, or failing its health check). | Check the endpoint's runtime.backends. |
NXDOMAIN for a name of another workspace |
The short name was searched in your own namespace. | Add the other workspace's namespace: <name>.<namespace>. |
astraeus.local names do not resolve at all |
The worker is on the host network. | Use the addresses in the environment, or run on the node network. |
| A name resolves but connections fail | A run's outbound policy (network.egress) or a firewall blocks the traffic. |
Check the run's Outbound network panel and network.egress.cluster. |
A run is refused at creation with NAME_NOT_DNS_SAFE |
The run, group or endpoint name cannot be a DNS name (only letters, digits, - and _; labels up to 63 characters, not starting or ending with -). |
Rename it. |