Topology#
Astralyx works out how your machines are connected, inside each machine and between them, from what each machine's agent reads. Placement uses it to keep a multi-machine run on as few switch hops as there is room for, and to tell you the network a run got and the all-reduce bandwidth to expect there. It also finds what looks miswired or degraded — a cable on the wrong rail's leaf, a link below its rate, a PCIe link below its width — and says where and how to fix it. Use this page to read your cluster's topology, to understand why a run got the network it got, to act on findings, to give the agent what it needs on a cloud, and to export the topology to Slurm.
Before you begin#
- Discovery needs nothing to start: every machine whose agent is recent reports its topology. A machine whose agent is older shows as not reporting it yet until its agent is updated (Update the agent).
- Two packages make it see more, when you install them on the machines
(the installer does not):
lldpd, for the switch each Ethernet port is cabled to, andinfiniband-diags, for the InfiniBand subnets (sminfo) and the fabric's map (ibnetdiscover). - Within a workspace you see only the machines of its pools, the switches they are cabled to, and the findings about them alone. On Astraeus Cloud, only your organisation's machines.
- In the API examples,
ASTRALYX_APIishttps://api.astralyx.cloud/v1andASTRALYX_TOKENan API token of the workspace (see REST API). The examples show a cluster of eight machines,gpu-01togpu-08, each with eight H100 GPUs and eight 400 Gb/s InfiniBand ports.
What is discovered#
Inside each machine#
| What | From |
|---|---|
| Each GPU's PCI address, NUMA node and PCIe link (generation and width, now and at most) | The GPU driver; sysfs (/sys/bus/pci/devices/…) |
| NVLinks up per GPU, of how many, and their bandwidth (25 GB/s each way per link; 50 GB/s on Blackwell) | The GPU driver, read with the machine's other GPU data |
| Multi-node NVLink fabric registration (GB200 NVL72): its state, and the NVLink domain once it completed | The GPU driver |
| Each RDMA port's device, netdev, link layer (InfiniBand, RoCE, EFA), state, rate, adapter, node GUID, PCI address, NUMA node and PCIe link | sysfs (/sys/class/infiniband/…) |
Each InfiniBand port's subnet: gid-<prefix> from its GID's subnet prefix when that is not the default, else sm-<guid> from its subnet manager |
sysfs; sminfo (from infiniband-diags) for the subnet manager |
How each GPU reaches each other GPU and each port: NV18 (18 NVLinks), PIX (one PCIe switch between them), PXB (several bridges), PHB (through a host bridge), NODE (another host bridge, same NUMA node), SYS (across NUMA nodes) — as nvidia-smi topo -m says it |
The PCI tree in sysfs |
The rails: the RDMA ports beside a GPU (PIX or PXB), numbered in their GPUs' order. Other RDMA ports (storage, front end) are not rails. On a GPU machine whose PCI tree cannot be read, its InfiniBand ports, in device order. |
The above |
| The front end: the fastest Ethernet link up that is not an RDMA port, and the switch its port is cabled to — the machine's top-of-rack switch | sysfs; LLDP (lldpctl, where lldpd runs) |
Between machines#
| What | From |
|---|---|
| The switch each Ethernet or RoCE port is cabled to, and its port there | LLDP: lldpctl -f json, where lldpd runs on the machine and the switches send LLDP |
| The fabric's map: every switch, every adapter and every cable of an InfiniBand subnet, with the rate each cable trained at | ibnetdiscover (from infiniband-diags), run by one machine per subnet, only where you allow it: see Sweeps |
On a cloud#
When the machine's firmware says which cloud it is on (or the agent is told
with --cloud), the agent reads the cloud's placement data:
| Cloud | Read from | Gives |
|---|---|---|
| AWS | DescribeInstanceTopology for the instance, signed with the instance's own role (it needs ec2:DescribeInstanceTopology) |
The network nodes above the instance, top down; its availability zone, which EFA does not leave (the fabric) |
| Google Cloud | The metadata server's physical_host |
The block and sub-block above the host, and the host. Google Cloud does not bound RDMA by block: no fabric |
| Azure | The instance metadata service (compute) |
The placement group (the InfiniBand fabric) |
| Oracle Cloud | The instance metadata service (rdmaTopologyData) |
The HPC island (the fabric), network block and local block, and the host |
| Nebius, Lambda, or any other | A file an integration keeps on the machine (--topology-file): see The topology file |
What the file says: the network path top down, the fabric, the rack |
On Azure, Nebius and Google Cloud, the RDMA ports are virtual functions:
the fabric's management does not answer them, so neither sminfo nor a
sweep works there. The cloud's data stands in.
What the agent does, and does not#
Discovery is light, rare, and changes nothing:
- The inside of the machine is read from sysfs with each machine snapshot (every 30 seconds): a few dozen small files. No tool is run for it.
- LLDP is read at most every 10 minutes.
lldpctlanswers from whatlldpdalready knows: nothing goes on the wire for it. - The cloud is asked at most once an hour (a virtual machine can move to
another host), each request with a 2-second timeout. The metadata services
are on the machine's link-local address; on AWS, one signed
DescribeInstanceTopologycall goes toec2.<region>.amazonaws.com. - The subnet manager (
sminfo) is asked once per InfiniBand port, and again only when the port's subnet manager changes; an unanswered question is asked again after 10 minutes. A port whose GID prefix names its subnet needs no question. - The fabric's map (
ibnetdiscover) is made only by machines you allow, only when asked: see Sweeps. - What it finds is sent only when it changes: a few kilobytes for an eight-GPU server. A change in a port's state or rate, a PCIe link or the NVLinks counts only once seen twice in a row, so a port flapping for seconds is not reported each time.
- Discovery only reads. It configures nothing on the machine, its ports or the switches.
Sweeps of an InfiniBand fabric#
A sweep maps a whole InfiniBand subnet: its switches, the hosts' adapters
and every cable, with the rate each trained at. It is what tells leaves
from spines, which leaf each port reaches, and which cable is on the wrong
rail. ibnetdiscover sends management datagrams to every switch and
adapter of the subnet, so:
- It runs only on machines whose agent runs with
--topology-sweep(see Agent settings), withibnetdiscoverinstalled and an InfiniBand port up that knows its subnet and is not a virtual function. Allow one or two machines per subnet: allowing more does not sweep more often, since one machine per subnet is asked. - One machine per subnet sweeps — the one that swept last, while it still may — when the subnet has no map yet, once a day, and when its machines or ports change (a machine joins or leaves, a port goes down or comes back). Never twice within 30 minutes. A sweep that runs longer than 120 seconds is stopped.
- Without a sweep, InfiniBand machines still get their fabric from their subnet; leaves, spines and cabling findings need a sweep, LLDP (RoCE and Ethernet) or a cloud's data.
Each subnet's last sweep is shown with the topology (in the console, under Sweeps on the Tiers tab): which machine made it, when, how many switches and cables it found, how many adapters belong to no machine shown, or why the last attempt failed. A workspace does not start sweeps: the fabrics are mapped on their own. After recabling, a port that stayed down for a minute or more changes the subnet's ports, so the subnet is mapped again, no sooner than 30 minutes after its last attempt; otherwise, at its daily sweep.
Fabrics, leaf groups and spine groups#
Astralyx merges every machine's discovery, the fabrics' maps and the clouds' data into one topology:
- Fabrics: machines that reach each other over RDMA. An InfiniBand
fabric is its subnet (
sm-<guid>,gid-<prefix>; machines whose rails are on several subnets get a name for the set,rails-<n>x-<hash>), or, without one, what the cloud bounds RDMA by:<cloud>-<id>, for example an AWS availability zone or an Azure placement group. Machines on two fabrics cannot talk over RDMA. - Tiers come from the cabling: tier 0 are the switches machines' ports reach (leaves); tier 1, the switches one hop above them (spines).
- Leaf groups. On a rail-optimised fabric each machine's rail k is
cabled to a leaf of rail k, and the leaves of every rail serve the same
machines. Leaves that share at least half of their machines are one
leaf group: its machines are one hop apart on every rail. A group is
named after its switches:
leaf-r0-7-g0forleaf-r0-g0…leaf-r7-g0. A machine's leaf group is the one most of its compute ports reach, so one cable plugged into the wrong leaf is a finding, not a new group. - Spine groups: leaf groups that share a spine (
spine-0-1forspine-0andspine-1). Machines under one spine group but different leaf groups are three hops apart. - Oversubscription: a leaf's down links over its up links, Gb/s over
Gb/s, the worst leaf of its group;
1is non-blocking. - NVLink domains: machines whose GPUs registered with one multi-node NVLink fabric (the trays of a GB200 NVL72). A machine joins its domain only once its GPUs completed registration.
- Cables between switches, from the fabrics' maps, each with the rate it trained at and its peers' rate.
flowchart TB
subgraph SG["spine group spine-0-1"]
S0[spine-0]
S1[spine-1]
end
subgraph G0["leaf group leaf-r0-7-g0"]
A0["leaf-r0-g0 (rail 0)"]
A7["… leaf-r7-g0 (rail 7)"]
end
subgraph G1["leaf group leaf-r0-7-g1"]
B0["leaf-r0-g1 (rail 0)"]
B7["… leaf-r7-g1 (rail 7)"]
end
S0 & S1 --- A0 & A7 & B0 & B7
A0 & A7 --- M0["gpu-01 … gpu-04"]
B0 & B7 --- M1["gpu-05 … gpu-08"]
Labels#
Discovery derives labels on each machine that reports its topology.
Placement reads them like any other label, so machine_selection.match_labels
can select them too ({"topology.astraeus.io/rails": "8"} keeps a run on
machines with all eight rails up).
| Label | What | Derived from | An admin can set it |
|---|---|---|---|
topology.astraeus.io/fabric |
Machines that reach each other over RDMA | The InfiniBand subnets (sm-<guid>, gid-<prefix>, rails-<n>x-<hash>); else the subnet a sweep mapped; else the cloud's data (<cloud>-<id>) |
Yes |
topology.astraeus.io/spine |
The spine group above the machine's leaf group | A sweep; else the cloud's network one level above the machine's own | No |
topology.astraeus.io/leaf |
The leaf group (or a single leaf) the machine's compute ports reach | A sweep or LLDP; else the cloud's lowest network node (<cloud>-<node>) |
No |
topology.astraeus.io/rack |
The rack | The top-of-rack switch LLDP sees on the machine's front-end port; else the cloud's rack | Yes |
topology.astraeus.io/rails |
How many compute rails are up | sysfs | No |
topology.astraeus.io/nvlink-domain |
The multi-node NVLink domain | The GPUs' NVLink fabric registration, once it completed | No |
- An admin's label always wins for placement. Organisation admins set
rackandfabric(andpoolandsite) in Where it is on the machine's page; see Pools, labels and topology. The derived value stays visible beside it, so a label that no longer matches the cabling shows. - Derived labels are on the machine's topology, not among its labels:
GET /machines/{name}does not list them, and they change by themselves when the cabling does. - A machine whose agent does not report its topology yet gets only its fabric, from its InfiniBand subnets.
- Values are letters, digits,
-,_and., at most 63 characters; a longer name keeps its end behind a short digest.
Findings#
Each finding says what it found, the evidence, and the fix. warning:
something is wrong. info: worth knowing.
| Kind | Severity | What it means |
|---|---|---|
disjoint-fabric |
info |
Fabrics of one kind (InfiniBand, RoCE, EFA) that do not reach each other. Nothing to do if they are meant to be separate: a run that needs RDMA is kept within one. |
disjoint-fabric |
warning |
Machines an admin labelled one fabric are on several that do not reach each other. A run kept within that label across them hangs at its first collective. |
rail-misaligned |
warning |
A compute port cabled to another rail's leaf, or to a leaf of another leaf group than the machine's other rails. That rail's traffic crosses the spines. |
link-degraded |
warning |
A port that trained below the rate the same adapter trains at elsewhere, or a cable between switches below the fabric's other cables between the same tiers. |
pcie-degraded |
warning |
A GPU's or an RDMA port's PCIe link narrower or slower than it can be (Gen5 x8 (of Gen5 x16)). |
ports-missing |
warning |
A machine with fewer compute ports than machines of the same kind (GPU model and count, cloud), or with some of them down. |
nvlink-incomplete |
warning |
GPUs that have not completed their NVLink fabric registration: the machine is in no NVLink domain, and a run that needs one is not placed there. |
provider-disagrees |
warning |
The cloud's placement data puts a machine under another leaf, or on another fabric, than its cabling or its subnets. Placement follows the cabling. |
A finding in the example cluster, as the API shows it:
{
"id": "rail/gpu-02/mlx5_3",
"kind": "rail-misaligned",
"severity": "warning",
"summary": "Miswired: gpu-02 mlx5_3 (rail 3) is cabled to leaf-r4-g0 port 5, rail 4's leaf",
"evidence": [
"the fabric's map (ibnetdiscover from gpu-01, 2026-10-05 03:12 UTC): leaf-r4-g0 carries rail 4: port 1: gpu-01 mlx5_4 (rail 4); port 5: gpu-02 mlx5_3 (rail 3); port 2: gpu-02 mlx5_4 (rail 4); port 3: gpu-03 mlx5_4 (rail 4); port 4: gpu-04 mlx5_4 (rail 4)",
"rail 3's traffic between gpu-02 and its leaf group's machines crosses the spines: about 85% of line rate on that rail"
],
"fix": "Move gpu-02 mlx5_3's cable from leaf-r4-g0 port 5 to a free port of leaf-r3-g0 (rail 3's leaf).",
"machines": ["gpu-02"],
"switches": ["ib:0xfc6a1c0300001040"],
"since": "2026-10-05T03:12:10.004417Z"
}
Other findings read the same way:
| Example summary | Fix it gives |
|---|---|
gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace |
Reseat or replace the cable and transceiver (mlxlink -d mlx5_5 -m shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out. |
gpu-03 GPU 5's PCIe link trained at Gen5 x8 (of Gen5 x16) |
Reseat the GPU or its riser (lspci -vv -s <address> shows LnkSta; AER errors in the kernel log say whether it retrained down): host-to-GPU copies run at that width. |
gpu-04: 7 of 8 compute ports up (mlx5_6 down) |
Check the cable, the transceiver and the switch port (ibstat mlx5_6). |
gpu-05 has 7 compute port(s); its peers 8: an adapter is missing (off the PCI bus, or without its driver) |
Check that every adapter is seated and has its driver (lspci, ibstat). |
tray-07: GPU(s) 0, 1, 2, 3 have not completed NVLink fabric registration (in progress): the machine is in no NVLink domain |
Check the NVLink switch trays, their fabric manager, and nvidia-smi -q (Fabric) on the machine. |
What a finding does:
- A warning sets the machine's condition
TopologyHealthytoFalse, with the finding's kind as the reason (RailMisaligned,LinkDegraded,PcieDegraded,PortsMissing,NvlinkDomainIncomplete,ProviderDisagrees,FabricDisjoint) and the summaries as its message. It is advisory: work is still placed on the machine. Its change is an event of the machine, and an alert rule onmachine_conditionis told (see Alerts). It turnsTrueagain (AsDiscovered) once nothing is wrong. - An info finding is an event on its machines when it is found
(
Topology: …) and when it is resolved (Topology: no longer so: …). - A finding keeps the time it was first found (
since) and anidthat stays the same for the same thing. - Checks quote it. A check whose measurement falls short where a finding explains why — a pair slow on a miswired rail, copies slow on a narrow PCIe link — names the finding and its fix, and the Cluster Report lists the findings beside each machine's checks (Acceptance and checks).
How placement uses it#
A run placed whole (a gang) climbs down this ladder and stops at the first rung where every worker fits:
- One machine.
- One multi-node NVLink domain.
- One leaf group, over RDMA of one kind: one hop on every rail.
- One InfiniBand fabric and one rack.
- One InfiniBand fabric.
- RDMA of one kind in one rack: InfiniBand whose fabric is not known, or RoCE or EFA within its fabric.
- The same across racks.
- One rack, over Ethernet.
-
Ethernet.
-
Never across fabrics that do not reach each other. Once any machine of the cluster has a fabric — derived or set by an admin — machines with RDMA are grouped by their fabric. A run that needs RDMA (
network.interconnect: rdmaorinfiniband, ormin_gbps) waits rather than span two; a run withinterconnect: autothat no fabric can hold is placed over Ethernet. - A run placed on one leaf group gets what a run on a fabric gets: the
machine's RDMA devices,
IPC_LOCKandNCCL_IB_HCAnaming the ports up and fast enough. Each worker'splacement_networksaysleaf:<leaf group>. - The rules a run can set (keep within, spread over, link speed, patience) are in GPUs and placement.
The expected all-reduce bandwidth#
Each run of several workers placed whole records the network its workers
got and the all-reduce bandwidth to expect there — bus bandwidth, as
nccl-tests reports it:
| Placed on | Expected bus bandwidth, GB/s |
|---|---|
| One machine, one NVLink domain | The slowest GPU's NVLink bandwidth × 0.92 |
| One leaf group | Rails × the slowest rail's rate (Gb/s ÷ 8) × 0.92 |
| Several leaf groups (a fabric, RDMA) | Rails × the slowest rail's rate (Gb/s ÷ 8) × 0.92 × 0.85, divided by the worst oversubscription of their leaves |
| A rack or anywhere over Ethernet | The slowest front-end link (Gb/s ÷ 8) × 0.7 × 0.92 |
It is never more than NVLink gives inside the machines, and it is an
estimate from the links' rates, not a measurement: check it with an
all-reduce (Check the network first),
or run the network checks, which measure all-reduce between your machines
and hold it to this estimate, at the rates the links are rated for
(Acceptance and checks).
For eight 400 Gb/s rails: one leaf switch expects about 368 GB/s, and
2 leaves under one spine group about 313 GB/s. One rail at 200 Gb/s
halves both: the slowest rail sets the pace.
A run that asked to wait for a better network (topology.patience_seconds)
says what it waits for, and both bandwidths:
Waiting up to 25 min more for one leaf switch (leaf-r0-7-g0), ~368 GB/s all-reduce (could start now on InfiniBand fabric sm-0x0002c90300a1b2c3, ~313 GB/s all-reduce)
To see the network a run got, and what to expect there:
Open the run. Asked and got shows the network its workers got and,
for a run placed whole, what to expect there (expect ~313 GB/s
all-reduce), with the estimate's words and basis.
astra does not show a run's network. Use the console or the API.
$ curl -sS "$ASTRALYX_API/runs/train" -H "Authorization: Bearer $ASTRALYX_TOKEN" | jq .network
{
"network": "fabric:sm-0x0002c90300a1b2c3",
"words": "InfiniBand fabric sm-0x0002c90300a1b2c3",
"machines": [
"gpu-03",
"gpu-07"
],
"estimate": {
"busbw_gbps": 312.8,
"words": "2 leaves under one spine group",
"basis": "8 × 400 Gb/s rails, through the spines (× 0.85), × 0.92",
"leaves": 2,
"spines": 1
},
"placed_at": "2026-10-05T09:20:11.418093Z"
}
network is present on runs of several workers placed whole, written
each time the run is placed.
See the topology#
Open Compute → Topology and choose the cluster. The page has three tabs:
- Racks: the machines drawn rack by rack. Rack and fabric are an admin's labels, else what discovery found; each machine shows its leaf group, when one is known.
- Tiers: Fabrics (kind, machines, leaf groups, what told them apart); Spines, leaves and machines, each spine group with its leaf groups, each listing its leaf switches by rail, its oversubscription when above 1:1 and how it was found, then the leaf groups with no spine group known and the machines under no leaf known; NVLink domains, with their GPUs and the trays not complete; and Sweeps.
- Findings: each finding's severity, summary and fix, its evidence (folded) and the machines it is about.
$ astra astraeus topology
Topology of 8 machine(s), merged 2026-10-05T09:12:41.208734512Z
Fabrics (machines that reach each other over RDMA)
FABRIC KIND MACHINES
sm-0x0002c90300a1b2c3 InfiniBand gpu-[01-08] found by sminfo
Tiers (spine > leaf > machines)
spine spine-0-1 (fabric sm-0x0002c90300a1b2c3)
leaf leaf-r0-7-g0 gpu-[01-04] (ibnetdiscover)
leaf leaf-r0-7-g1 gpu-[05-08] (ibnetdiscover)
Findings (2)
! gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace
fix: Reseat or replace the cable and transceiver of gpu-06 mlx5_5 (`mlxlink -d mlx5_5 -m` shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.
! Miswired: gpu-02 mlx5_3 (rail 3) is cabled to leaf-r4-g0 port 5, rail 4's leaf
fix: Move gpu-02 mlx5_3's cable from leaf-r4-g0 port 5 to a free port of leaf-r3-g0 (rail 3's leaf).
Sweeps of the InfiniBand fabrics
sm-0x0002c90300a1b2c3 by gpu-01 6h: 18 switches, 192 cables
! is a warning, i an info finding. Machines whose agent does not
report its topology yet are listed last. --json prints the whole
topology as the API answers it.
$ curl -sS "$ASTRALYX_API/topology" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
| jq -c '.groups[] | {id, tier, parent, machines}'
{"id":"leaf-r0-7-g0","tier":0,"parent":"spine-0-1","machines":["gpu-01","gpu-02","gpu-03","gpu-04"]}
{"id":"leaf-r0-7-g1","tier":0,"parent":"spine-0-1","machines":["gpu-05","gpu-06","gpu-07","gpu-08"]}
{"id":"spine-0-1","tier":1,"parent":null,"machines":["gpu-01","gpu-02","gpu-03","gpu-04","gpu-05","gpu-06","gpu-07","gpu-08"]}
The answer has fabrics, groups, switches, links (cables
between switches), machines, domains, findings and sweeps;
every field is in the
API reference.
Your AI assistant can read it too: the tools get_topology and
get_machine_topology (Connect your AI assistant).
See one machine's topology#
Open Compute → Machines and select the machine. Two sections show its topology:
- Its place on the network: the findings about it; then, for its fabric, spine group, leaf group, rack and rails up, what discovery Found, where From, the admin's Label, and what Placement uses; its NVLink domain; its top-of-rack switch, front-end speed and the cloud's data; and, folded, how it was found: what each source answered.
- Inside the machine: the GPU × GPU and GPU × port matrix, with each GPU's NUMA node, PCIe link and NVLinks; and its ports: rail, state, rate against what it should train at, PCIe, NUMA, nearest GPU, the switch and port it is Cabled to, and who saw that.
$ astra astraeus topology --machine gpu-02
gpu-02 fabric sm-0x0002c90300a1b2c3 spine spine-0-1 leaf leaf-r0-7-g0 rack r12
Derived: fabric=sm-0x0002c90300a1b2c3 (sminfo), leaf=leaf-r0-7-g0 (ibnetdiscover), rack=tor-r12 (lldp), rails=8 (sysfs), spine=spine-0-1 (ibnetdiscover)
Set by an admin (these win): rack=r12
GPUs and ports (NV#: NVLinks; PIX: one PCIe switch; PXB: several bridges; PHB: the host bridge; NODE: same NUMA node; SYS: across)
GPU0 GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7 mlx5_0 mlx5_1 mlx5_2 mlx5_3 mlx5_4 mlx5_5 mlx5_6 mlx5_7 NUMA PCIe
GPU0 X NV18 NV18 NV18 NV18 NV18 NV18 NV18 PIX NODE NODE NODE SYS SYS SYS SYS 0 Gen5 x16
GPU1 NV18 X NV18 NV18 NV18 NV18 NV18 NV18 NODE PIX NODE NODE SYS SYS SYS SYS 0 Gen5 x16
GPU2 NV18 NV18 X NV18 NV18 NV18 NV18 NV18 NODE NODE PIX NODE SYS SYS SYS SYS 0 Gen5 x16
GPU3 NV18 NV18 NV18 X NV18 NV18 NV18 NV18 NODE NODE NODE PIX SYS SYS SYS SYS 0 Gen5 x16
GPU4 NV18 NV18 NV18 NV18 X NV18 NV18 NV18 SYS SYS SYS SYS PIX NODE NODE NODE 1 Gen5 x16
GPU5 NV18 NV18 NV18 NV18 NV18 X NV18 NV18 SYS SYS SYS SYS NODE PIX NODE NODE 1 Gen5 x16
GPU6 NV18 NV18 NV18 NV18 NV18 NV18 X NV18 SYS SYS SYS SYS NODE NODE PIX NODE 1 Gen5 x16
GPU7 NV18 NV18 NV18 NV18 NV18 NV18 NV18 X SYS SYS SYS SYS NODE NODE NODE PIX 1 Gen5 x16
Ports
PORT RAIL STATE RATE SWITCH SWITCH PORT SEEN BY
mlx5_0/1 0 up 400 Gb/s leaf-r0-g0 2 ibnetdiscover
mlx5_1/1 1 up 400 Gb/s leaf-r1-g0 2 ibnetdiscover
mlx5_2/1 2 up 400 Gb/s leaf-r2-g0 2 ibnetdiscover
mlx5_3/1 3 up 400 Gb/s leaf-r4-g0 5 ibnetdiscover
mlx5_4/1 4 up 400 Gb/s leaf-r4-g0 2 ibnetdiscover
mlx5_5/1 5 up 400 Gb/s leaf-r5-g0 2 ibnetdiscover
mlx5_6/1 6 up 400 Gb/s leaf-r6-g0 2 ibnetdiscover
mlx5_7/1 7 up 400 Gb/s leaf-r7-g0 2 ibnetdiscover
Findings
! Miswired: gpu-02 mlx5_3 (rail 3) is cabled to leaf-r4-g0 port 5, rail 4's leaf
fix: Move gpu-02 mlx5_3's cable from leaf-r4-g0 port 5 to a free port of leaf-r3-g0 (rail 3's leaf).
A port below its rate shows both (200 Gb/s (of 400)); a PCIe link
below its width, both (Gen5 x8 (of Gen5 x16)). A machine on a cloud
also shows the cloud's network path and fabric.
$ curl -sS "$ASTRALYX_API/machines/gpu-02/topology" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
| jq -c '.reported.sources[]'
{"name":"sysfs","ok":true,"detail":"8 GPU(s), 8 RDMA port(s), 8 rail(s)"}
{"name":"lldp","ok":true,"detail":"1 neighbour(s)"}
{"name":"sminfo","ok":true,"detail":"8 of 8 InfiniBand port(s) know their subnet"}
| Field | What |
|---|---|
reported |
What the agent last read: GPUs, RDMA ports, the GPU×GPU and GPU×port matrices (gpu_gpu, gpu_nic), the top-of-rack switch (tor), the cloud's data (provider), and what each source said (sources). null for an agent that does not report its topology yet. |
machine |
Its place in the merged topology: fabric, spine and leaf groups, rack, NVLink domain, ports and where their cables go, labels (derived, admin, sources). |
placement |
What placement takes from it: the derived labels, its rails' count and rate, its leaf group's oversubscription, its NVLink and front-end rates. |
findings |
The findings about it. |
A machine outside the workspace's pools answers 404 NODE_NOT_FOUND.
Export to Slurm#
The topology exports as Slurm's topology.conf, for a Slurm cluster on the
same machines — for example machines installed in
observe-only mode beside the Slurm cluster that runs
them:
| Form | Slurm plugin | What it holds |
|---|---|---|
tree (the default) |
topology/tree |
A switch per leaf group with its machines (Nodes=), a switch per spine group with its leaf groups (Switches=), and a switch named after the fabric above them where a fabric has several spine groups, or leaf groups and no spine known. Machines on a fabric under no known leaf are under <fabric>-unplaced. Fabrics that do not reach each other are separate trees. |
block |
topology/block |
A block per NVLink domain (nvl-<domain>), and, for machines in none, a block per leaf group. |
Machines in no switch or block are listed in a closing comment. Machine
names are Astralyx's: the hosts' names, unless a machine was installed with
another --name. Slurm's node names must match them. Within a workspace,
the export holds only the machines of its pools.
- Open Compute → Topology and select Export for Slurm.
- Choose Tree or Block.
- Select Download topology.conf, or copy it.
$ astra astraeus topology --slurm tree
# topology.conf for Slurm's topology/tree, written by Astralyx from the cluster's discovered topology (2026-10-05 09:12 UTC).
# Leaf switches are leaf groups: on a rail-optimised fabric, the leaves of every rail that serve the same machines.
SwitchName=leaf-r0-7-g0 Nodes=gpu-[01-04]
SwitchName=leaf-r0-7-g1 Nodes=gpu-[05-08]
SwitchName=spine-0-1 Switches=leaf-r0-7-g0,leaf-r0-7-g1
--slurm alone is --slurm tree; --slurm block writes blocks.
$ curl -sS "$ASTRALYX_API/topology/slurm?form=block" -H "Authorization: Bearer $ASTRALYX_TOKEN"
# topology.conf for Slurm's topology/block, written by Astralyx from the cluster's discovered topology (2026-10-05 09:12 UTC).
# Blocks are NVLink domains where machines are in one, else leaf groups (one hop on every rail).
BlockName=leaf-r0-7-g0 Nodes=gpu-[01-04]
BlockName=leaf-r0-7-g1 Nodes=gpu-[05-08]
The answer is text/plain. Another form answers 400 INVALID_FORM.
Save it as Slurm's topology.conf and name the plugin in slurm.conf
(TopologyPlugin=topology/tree, or topology/block).
Agent settings#
Discovery needs no setting. Three tell the machine's agent more:
| Flag | Variable | Default | Description |
|---|---|---|---|
--topology-sweep |
ASTRAEUS_TOPOLOGY_SWEEP |
false |
Let this machine map its InfiniBand fabric with ibnetdiscover when asked: one machine per subnet, about once a day. See Sweeps. |
--cloud <name> |
ASTRAEUS_CLOUD |
What the firmware says | The cloud this machine is on, when its firmware does not say: aws, gcp, azure, oci, nebius or lambda. Its placement data is read from that cloud's metadata service (Nebius and Lambda: from --topology-file). |
--topology-file <path> |
ASTRAEUS_TOPOLOGY_FILE |
none | The cloud's placement data for this machine, as a file an integration keeps current, read instead of the metadata service, again every hour. See The topology file. |
The installer does not write them. Set them in a systemd drop-in, which survives reruns of the installer, then restart the agent; its workers keep running:
sudo mkdir -p /etc/systemd/system/astraeus-agent.service.d
sudo tee /etc/systemd/system/astraeus-agent.service.d/topology.conf >/dev/null <<'EOF'
[Service]
Environment=ASTRAEUS_TOPOLOGY_SWEEP=true
EOF
sudo systemctl daemon-reload
sudo systemctl restart astraeus-agent
On AWS
The instance's role needs ec2:DescribeInstanceTopology, and the
machine must reach ec2.<region>.amazonaws.com on 443 directly: that
call does not go through HTTPS_PROXY. Without them, the machine's
provider source says why, and the machine gets no leaf or spine from
AWS.
The topology file#
--topology-file names a JSON file with the cloud's placement data for the
machine, in this shape:
| Field | Type | Description |
|---|---|---|
provider |
string | The cloud (nebius, lambda, any name). Empty: --cloud, or the cloud the firmware says, else operator. Derived names start with it: <provider>-<value>. |
instance |
string | The instance's id. |
path |
list of strings | The network above the machine, from the top down to its own switch or block. The last is its leaf (topology.astraeus.io/leaf), the one before it its spine. |
fabric |
string | Where RDMA reaches: machines with the same fabric reach each other (topology.astraeus.io/fabric, when the InfiniBand subnets do not say). |
rack |
string | The rack (topology.astraeus.io/rack, when LLDP does not say). |
host |
string | The physical host. |
On Nebius, write the instance's InfiniBand topology path as path and its
GPU cluster as fabric; on Lambda, the cluster as fabric, the rack as
rack, and the rail group as path. Below, the values stand for your
cloud's:
{
"provider": "nebius",
"instance": "<instance id>",
"path": ["<fabric>", "<spine group>", "<leaf>"],
"fabric": "<GPU cluster id>"
}
Name it in the agent's drop-in
(Environment=ASTRAEUS_TOPOLOGY_FILE=/etc/astraeus/topology.json, as in
Agent settings). The agent reads it again every hour; a
file it cannot read or parse is the machine's provider source's error.
Reference: the API#
| Method and path | What | Role |
|---|---|---|
GET /topology |
The topology of the machines the workspace may use | viewer or above |
GET /topology/slurm?form=tree\|block |
Slurm's topology.conf, as text |
viewer or above |
GET /machines/{name}/topology |
One machine's topology | viewer or above |
GET /runs/{name} |
The run, with network: where a gang went and the bandwidth to expect |
viewer or above |
Every field is in the API reference. Fabrics are mapped again on their own (see Sweeps): a workspace cannot ask for a sweep.
Limits#
| Limit | Value |
|---|---|
| GPUs a machine's topology may list | 64 |
| RDMA ports a machine's topology may list | 256 |
| A sweep's duration | 120 s, then stopped |
| Sweeps of one subnet | One machine at a time; at least 30 minutes apart; again after 24 hours unchanged |
| LLDP read again | Every 10 minutes |
| The cloud's data read again | Every hour |
| A metadata request | 2 s |
A machine's report over these limits is refused.
Troubleshooting#
Each machine's reported.sources (GET /machines/{name}/topology) says
what each source found, or why it could not.
| Symptom | Cause | Fix |
|---|---|---|
| A machine is listed as not reporting its topology yet | Its agent is older than topology discovery. | Update its agent (Update the agent). |
| InfiniBand machines have a fabric but no leaf or spine | No sweep: no machine of the subnet may sweep, or it cannot. | Allow one or two machines per subnet (--topology-sweep) with infiniband-diags installed. |
ibnetdiscover source: not installed (infiniband-diags) |
The machine may sweep, but the package is missing. | Install infiniband-diags on that machine. |
No ibnetdiscover source on the machine you allowed |
The setting did not reach the agent, or no InfiniBand port is up there. | Check the drop-in, sudo systemctl daemon-reload, restart the agent; check ibstat. |
The sweep says ibnetdiscover did not finish within 120s |
A very large fabric, or one that does not answer. | Check the fabric with ibnetdiscover by hand on that machine. |
ibnetdiscover source: no port the fabric answers (virtual functions, or no subnet manager) |
Virtual functions (Azure, Nebius, Google Cloud), or the subnet manager does not answer. | On those clouds, the cloud's data stands in; elsewhere, check the subnet manager. |
sminfo source: the subnet manager did not answer (sminfo) |
infiniband-diags is missing, or the subnet manager does not answer. |
Install infiniband-diags; check sminfo on the machine. The question is asked again after 10 minutes. |
lldp source: lldpctl is not installed (lldpd) |
No lldpd. |
Install and start lldpd; the switches must send LLDP. Without it, set the rack in Where it is. |
provider source: DescribeInstanceTopology: UnauthorizedOperation (the instance role needs ec2:DescribeInstanceTopology) |
The instance's role lacks the permission. | Add ec2:DescribeInstanceTopology to the role. |
provider source: the instance has no role: DescribeInstanceTopology needs one |
The instance has no role. | Attach a role with ec2:DescribeInstanceTopology. |
provider source: DescribeInstanceTopology does not list i-… |
AWS gives no topology for this instance type. | Nothing to do: the machine gets no leaf, spine or fabric from AWS. |
provider source: nebius's topology is not readable from the machine: leave it in a file (--topology-file) |
Nebius and Lambda are read from a file. | Keep a topology file on the machine. |
| A rack label differs from the derived one | An admin's label wins; the cabling says otherwise. | Fix the label in Where it is, or empty it to use the derived rack. |
| An all-reduce is well below the run's expected bandwidth | A degraded or miswired link on one of its machines, or the image lacks the RDMA libraries. | Read the findings about its machines; check NCCL_DEBUG=INFO (Multi-machine runs). |