Skip to content

Topology#

Astralyx works out how your machines are connected, inside each machine and between them, from what each machine's agent reads. Placement uses it to keep a multi-machine run on as few switch hops as there is room for, and to tell you the network a run got and the all-reduce bandwidth to expect there. It also finds what looks miswired or degraded — a cable on the wrong rail's leaf, a link below its rate, a PCIe link below its width — and says where and how to fix it. Use this page to read your cluster's topology, to understand why a run got the network it got, to act on findings, to give the agent what it needs on a cloud, and to export the topology to Slurm.

Before you begin#

  • Discovery needs nothing to start: every machine whose agent is recent reports its topology. A machine whose agent is older shows as not reporting it yet until its agent is updated (Update the agent).
  • Two packages make it see more, when you install them on the machines (the installer does not): lldpd, for the switch each Ethernet port is cabled to, and infiniband-diags, for the InfiniBand subnets (sminfo) and the fabric's map (ibnetdiscover).
  • Within a workspace you see only the machines of its pools, the switches they are cabled to, and the findings about them alone. On Astraeus Cloud, only your organisation's machines.
  • In the API examples, ASTRALYX_API is https://api.astralyx.cloud/v1 and ASTRALYX_TOKEN an API token of the workspace (see REST API). The examples show a cluster of eight machines, gpu-01 to gpu-08, each with eight H100 GPUs and eight 400 Gb/s InfiniBand ports.

What is discovered#

Inside each machine#

What From
Each GPU's PCI address, NUMA node and PCIe link (generation and width, now and at most) The GPU driver; sysfs (/sys/bus/pci/devices/…)
NVLinks up per GPU, of how many, and their bandwidth (25 GB/s each way per link; 50 GB/s on Blackwell) The GPU driver, read with the machine's other GPU data
Multi-node NVLink fabric registration (GB200 NVL72): its state, and the NVLink domain once it completed The GPU driver
Each RDMA port's device, netdev, link layer (InfiniBand, RoCE, EFA), state, rate, adapter, node GUID, PCI address, NUMA node and PCIe link sysfs (/sys/class/infiniband/…)
Each InfiniBand port's subnet: gid-<prefix> from its GID's subnet prefix when that is not the default, else sm-<guid> from its subnet manager sysfs; sminfo (from infiniband-diags) for the subnet manager
How each GPU reaches each other GPU and each port: NV18 (18 NVLinks), PIX (one PCIe switch between them), PXB (several bridges), PHB (through a host bridge), NODE (another host bridge, same NUMA node), SYS (across NUMA nodes) — as nvidia-smi topo -m says it The PCI tree in sysfs
The rails: the RDMA ports beside a GPU (PIX or PXB), numbered in their GPUs' order. Other RDMA ports (storage, front end) are not rails. On a GPU machine whose PCI tree cannot be read, its InfiniBand ports, in device order. The above
The front end: the fastest Ethernet link up that is not an RDMA port, and the switch its port is cabled to — the machine's top-of-rack switch sysfs; LLDP (lldpctl, where lldpd runs)

Between machines#

What From
The switch each Ethernet or RoCE port is cabled to, and its port there LLDP: lldpctl -f json, where lldpd runs on the machine and the switches send LLDP
The fabric's map: every switch, every adapter and every cable of an InfiniBand subnet, with the rate each cable trained at ibnetdiscover (from infiniband-diags), run by one machine per subnet, only where you allow it: see Sweeps

On a cloud#

When the machine's firmware says which cloud it is on (or the agent is told with --cloud), the agent reads the cloud's placement data:

Cloud Read from Gives
AWS DescribeInstanceTopology for the instance, signed with the instance's own role (it needs ec2:DescribeInstanceTopology) The network nodes above the instance, top down; its availability zone, which EFA does not leave (the fabric)
Google Cloud The metadata server's physical_host The block and sub-block above the host, and the host. Google Cloud does not bound RDMA by block: no fabric
Azure The instance metadata service (compute) The placement group (the InfiniBand fabric)
Oracle Cloud The instance metadata service (rdmaTopologyData) The HPC island (the fabric), network block and local block, and the host
Nebius, Lambda, or any other A file an integration keeps on the machine (--topology-file): see The topology file What the file says: the network path top down, the fabric, the rack

On Azure, Nebius and Google Cloud, the RDMA ports are virtual functions: the fabric's management does not answer them, so neither sminfo nor a sweep works there. The cloud's data stands in.

What the agent does, and does not#

Discovery is light, rare, and changes nothing:

  • The inside of the machine is read from sysfs with each machine snapshot (every 30 seconds): a few dozen small files. No tool is run for it.
  • LLDP is read at most every 10 minutes. lldpctl answers from what lldpd already knows: nothing goes on the wire for it.
  • The cloud is asked at most once an hour (a virtual machine can move to another host), each request with a 2-second timeout. The metadata services are on the machine's link-local address; on AWS, one signed DescribeInstanceTopology call goes to ec2.<region>.amazonaws.com.
  • The subnet manager (sminfo) is asked once per InfiniBand port, and again only when the port's subnet manager changes; an unanswered question is asked again after 10 minutes. A port whose GID prefix names its subnet needs no question.
  • The fabric's map (ibnetdiscover) is made only by machines you allow, only when asked: see Sweeps.
  • What it finds is sent only when it changes: a few kilobytes for an eight-GPU server. A change in a port's state or rate, a PCIe link or the NVLinks counts only once seen twice in a row, so a port flapping for seconds is not reported each time.
  • Discovery only reads. It configures nothing on the machine, its ports or the switches.

Sweeps of an InfiniBand fabric#

A sweep maps a whole InfiniBand subnet: its switches, the hosts' adapters and every cable, with the rate each trained at. It is what tells leaves from spines, which leaf each port reaches, and which cable is on the wrong rail. ibnetdiscover sends management datagrams to every switch and adapter of the subnet, so:

  • It runs only on machines whose agent runs with --topology-sweep (see Agent settings), with ibnetdiscover installed and an InfiniBand port up that knows its subnet and is not a virtual function. Allow one or two machines per subnet: allowing more does not sweep more often, since one machine per subnet is asked.
  • One machine per subnet sweeps — the one that swept last, while it still may — when the subnet has no map yet, once a day, and when its machines or ports change (a machine joins or leaves, a port goes down or comes back). Never twice within 30 minutes. A sweep that runs longer than 120 seconds is stopped.
  • Without a sweep, InfiniBand machines still get their fabric from their subnet; leaves, spines and cabling findings need a sweep, LLDP (RoCE and Ethernet) or a cloud's data.

Each subnet's last sweep is shown with the topology (in the console, under Sweeps on the Tiers tab): which machine made it, when, how many switches and cables it found, how many adapters belong to no machine shown, or why the last attempt failed. A workspace does not start sweeps: the fabrics are mapped on their own. After recabling, a port that stayed down for a minute or more changes the subnet's ports, so the subnet is mapped again, no sooner than 30 minutes after its last attempt; otherwise, at its daily sweep.

Fabrics, leaf groups and spine groups#

Astralyx merges every machine's discovery, the fabrics' maps and the clouds' data into one topology:

  • Fabrics: machines that reach each other over RDMA. An InfiniBand fabric is its subnet (sm-<guid>, gid-<prefix>; machines whose rails are on several subnets get a name for the set, rails-<n>x-<hash>), or, without one, what the cloud bounds RDMA by: <cloud>-<id>, for example an AWS availability zone or an Azure placement group. Machines on two fabrics cannot talk over RDMA.
  • Tiers come from the cabling: tier 0 are the switches machines' ports reach (leaves); tier 1, the switches one hop above them (spines).
  • Leaf groups. On a rail-optimised fabric each machine's rail k is cabled to a leaf of rail k, and the leaves of every rail serve the same machines. Leaves that share at least half of their machines are one leaf group: its machines are one hop apart on every rail. A group is named after its switches: leaf-r0-7-g0 for leaf-r0-g0 … leaf-r7-g0. A machine's leaf group is the one most of its compute ports reach, so one cable plugged into the wrong leaf is a finding, not a new group.
  • Spine groups: leaf groups that share a spine (spine-0-1 for spine-0 and spine-1). Machines under one spine group but different leaf groups are three hops apart.
  • Oversubscription: a leaf's down links over its up links, Gb/s over Gb/s, the worst leaf of its group; 1 is non-blocking.
  • NVLink domains: machines whose GPUs registered with one multi-node NVLink fabric (the trays of a GB200 NVL72). A machine joins its domain only once its GPUs completed registration.
  • Cables between switches, from the fabrics' maps, each with the rate it trained at and its peers' rate.
flowchart TB
    subgraph SG["spine group spine-0-1"]
        S0[spine-0]
        S1[spine-1]
    end
    subgraph G0["leaf group leaf-r0-7-g0"]
        A0["leaf-r0-g0 (rail 0)"]
        A7["… leaf-r7-g0 (rail 7)"]
    end
    subgraph G1["leaf group leaf-r0-7-g1"]
        B0["leaf-r0-g1 (rail 0)"]
        B7["… leaf-r7-g1 (rail 7)"]
    end
    S0 & S1 --- A0 & A7 & B0 & B7
    A0 & A7 --- M0["gpu-01 … gpu-04"]
    B0 & B7 --- M1["gpu-05 … gpu-08"]

Labels#

Discovery derives labels on each machine that reports its topology. Placement reads them like any other label, so machine_selection.match_labels can select them too ({"topology.astraeus.io/rails": "8"} keeps a run on machines with all eight rails up).

Label What Derived from An admin can set it
topology.astraeus.io/fabric Machines that reach each other over RDMA The InfiniBand subnets (sm-<guid>, gid-<prefix>, rails-<n>x-<hash>); else the subnet a sweep mapped; else the cloud's data (<cloud>-<id>) Yes
topology.astraeus.io/spine The spine group above the machine's leaf group A sweep; else the cloud's network one level above the machine's own No
topology.astraeus.io/leaf The leaf group (or a single leaf) the machine's compute ports reach A sweep or LLDP; else the cloud's lowest network node (<cloud>-<node>) No
topology.astraeus.io/rack The rack The top-of-rack switch LLDP sees on the machine's front-end port; else the cloud's rack Yes
topology.astraeus.io/rails How many compute rails are up sysfs No
topology.astraeus.io/nvlink-domain The multi-node NVLink domain The GPUs' NVLink fabric registration, once it completed No
  • An admin's label always wins for placement. Organisation admins set rack and fabric (and pool and site) in Where it is on the machine's page; see Pools, labels and topology. The derived value stays visible beside it, so a label that no longer matches the cabling shows.
  • Derived labels are on the machine's topology, not among its labels: GET /machines/{name} does not list them, and they change by themselves when the cabling does.
  • A machine whose agent does not report its topology yet gets only its fabric, from its InfiniBand subnets.
  • Values are letters, digits, -, _ and ., at most 63 characters; a longer name keeps its end behind a short digest.

Findings#

Each finding says what it found, the evidence, and the fix. warning: something is wrong. info: worth knowing.

Kind Severity What it means
disjoint-fabric info Fabrics of one kind (InfiniBand, RoCE, EFA) that do not reach each other. Nothing to do if they are meant to be separate: a run that needs RDMA is kept within one.
disjoint-fabric warning Machines an admin labelled one fabric are on several that do not reach each other. A run kept within that label across them hangs at its first collective.
rail-misaligned warning A compute port cabled to another rail's leaf, or to a leaf of another leaf group than the machine's other rails. That rail's traffic crosses the spines.
link-degraded warning A port that trained below the rate the same adapter trains at elsewhere, or a cable between switches below the fabric's other cables between the same tiers.
pcie-degraded warning A GPU's or an RDMA port's PCIe link narrower or slower than it can be (Gen5 x8 (of Gen5 x16)).
ports-missing warning A machine with fewer compute ports than machines of the same kind (GPU model and count, cloud), or with some of them down.
nvlink-incomplete warning GPUs that have not completed their NVLink fabric registration: the machine is in no NVLink domain, and a run that needs one is not placed there.
provider-disagrees warning The cloud's placement data puts a machine under another leaf, or on another fabric, than its cabling or its subnets. Placement follows the cabling.

A finding in the example cluster, as the API shows it:

{
  "id": "rail/gpu-02/mlx5_3",
  "kind": "rail-misaligned",
  "severity": "warning",
  "summary": "Miswired: gpu-02 mlx5_3 (rail 3) is cabled to leaf-r4-g0 port 5, rail 4's leaf",
  "evidence": [
    "the fabric's map (ibnetdiscover from gpu-01, 2026-10-05 03:12 UTC): leaf-r4-g0 carries rail 4: port 1: gpu-01 mlx5_4 (rail 4); port 5: gpu-02 mlx5_3 (rail 3); port 2: gpu-02 mlx5_4 (rail 4); port 3: gpu-03 mlx5_4 (rail 4); port 4: gpu-04 mlx5_4 (rail 4)",
    "rail 3's traffic between gpu-02 and its leaf group's machines crosses the spines: about 85% of line rate on that rail"
  ],
  "fix": "Move gpu-02 mlx5_3's cable from leaf-r4-g0 port 5 to a free port of leaf-r3-g0 (rail 3's leaf).",
  "machines": ["gpu-02"],
  "switches": ["ib:0xfc6a1c0300001040"],
  "since": "2026-10-05T03:12:10.004417Z"
}

Other findings read the same way:

Example summary Fix it gives
gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace Reseat or replace the cable and transceiver (mlxlink -d mlx5_5 -m shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.
gpu-03 GPU 5's PCIe link trained at Gen5 x8 (of Gen5 x16) Reseat the GPU or its riser (lspci -vv -s <address> shows LnkSta; AER errors in the kernel log say whether it retrained down): host-to-GPU copies run at that width.
gpu-04: 7 of 8 compute ports up (mlx5_6 down) Check the cable, the transceiver and the switch port (ibstat mlx5_6).
gpu-05 has 7 compute port(s); its peers 8: an adapter is missing (off the PCI bus, or without its driver) Check that every adapter is seated and has its driver (lspci, ibstat).
tray-07: GPU(s) 0, 1, 2, 3 have not completed NVLink fabric registration (in progress): the machine is in no NVLink domain Check the NVLink switch trays, their fabric manager, and nvidia-smi -q (Fabric) on the machine.

What a finding does:

  • A warning sets the machine's condition TopologyHealthy to False, with the finding's kind as the reason (RailMisaligned, LinkDegraded, PcieDegraded, PortsMissing, NvlinkDomainIncomplete, ProviderDisagrees, FabricDisjoint) and the summaries as its message. It is advisory: work is still placed on the machine. Its change is an event of the machine, and an alert rule on machine_condition is told (see Alerts). It turns True again (AsDiscovered) once nothing is wrong.
  • An info finding is an event on its machines when it is found (Topology: …) and when it is resolved (Topology: no longer so: …).
  • A finding keeps the time it was first found (since) and an id that stays the same for the same thing.
  • Checks quote it. A check whose measurement falls short where a finding explains why — a pair slow on a miswired rail, copies slow on a narrow PCIe link — names the finding and its fix, and the Cluster Report lists the findings beside each machine's checks (Acceptance and checks).

How placement uses it#

A run placed whole (a gang) climbs down this ladder and stops at the first rung where every worker fits:

  1. One machine.
  2. One multi-node NVLink domain.
  3. One leaf group, over RDMA of one kind: one hop on every rail.
  4. One InfiniBand fabric and one rack.
  5. One InfiniBand fabric.
  6. RDMA of one kind in one rack: InfiniBand whose fabric is not known, or RoCE or EFA within its fabric.
  7. The same across racks.
  8. One rack, over Ethernet.
  9. Ethernet.

  10. Never across fabrics that do not reach each other. Once any machine of the cluster has a fabric — derived or set by an admin — machines with RDMA are grouped by their fabric. A run that needs RDMA (network.interconnect: rdma or infiniband, or min_gbps) waits rather than span two; a run with interconnect: auto that no fabric can hold is placed over Ethernet.

  11. A run placed on one leaf group gets what a run on a fabric gets: the machine's RDMA devices, IPC_LOCK and NCCL_IB_HCA naming the ports up and fast enough. Each worker's placement_network says leaf:<leaf group>.
  12. The rules a run can set (keep within, spread over, link speed, patience) are in GPUs and placement.

The expected all-reduce bandwidth#

Each run of several workers placed whole records the network its workers got and the all-reduce bandwidth to expect there — bus bandwidth, as nccl-tests reports it:

Placed on Expected bus bandwidth, GB/s
One machine, one NVLink domain The slowest GPU's NVLink bandwidth × 0.92
One leaf group Rails × the slowest rail's rate (Gb/s ÷ 8) × 0.92
Several leaf groups (a fabric, RDMA) Rails × the slowest rail's rate (Gb/s ÷ 8) × 0.92 × 0.85, divided by the worst oversubscription of their leaves
A rack or anywhere over Ethernet The slowest front-end link (Gb/s ÷ 8) × 0.7 × 0.92

It is never more than NVLink gives inside the machines, and it is an estimate from the links' rates, not a measurement: check it with an all-reduce (Check the network first), or run the network checks, which measure all-reduce between your machines and hold it to this estimate, at the rates the links are rated for (Acceptance and checks). For eight 400 Gb/s rails: one leaf switch expects about 368 GB/s, and 2 leaves under one spine group about 313 GB/s. One rail at 200 Gb/s halves both: the slowest rail sets the pace.

A run that asked to wait for a better network (topology.patience_seconds) says what it waits for, and both bandwidths:

Waiting up to 25 min more for one leaf switch (leaf-r0-7-g0), ~368 GB/s all-reduce (could start now on InfiniBand fabric sm-0x0002c90300a1b2c3, ~313 GB/s all-reduce)

To see the network a run got, and what to expect there:

Open the run. Asked and got shows the network its workers got and, for a run placed whole, what to expect there (expect ~313 GB/s all-reduce), with the estimate's words and basis.

astra does not show a run's network. Use the console or the API.

$ curl -sS "$ASTRALYX_API/runs/train" -H "Authorization: Bearer $ASTRALYX_TOKEN" | jq .network
{
  "network": "fabric:sm-0x0002c90300a1b2c3",
  "words": "InfiniBand fabric sm-0x0002c90300a1b2c3",
  "machines": [
    "gpu-03",
    "gpu-07"
  ],
  "estimate": {
    "busbw_gbps": 312.8,
    "words": "2 leaves under one spine group",
    "basis": "8 × 400 Gb/s rails, through the spines (× 0.85), × 0.92",
    "leaves": 2,
    "spines": 1
  },
  "placed_at": "2026-10-05T09:20:11.418093Z"
}

network is present on runs of several workers placed whole, written each time the run is placed.

See the topology#

Open Compute → Topology and choose the cluster. The page has three tabs:

  • Racks: the machines drawn rack by rack. Rack and fabric are an admin's labels, else what discovery found; each machine shows its leaf group, when one is known.
  • Tiers: Fabrics (kind, machines, leaf groups, what told them apart); Spines, leaves and machines, each spine group with its leaf groups, each listing its leaf switches by rail, its oversubscription when above 1:1 and how it was found, then the leaf groups with no spine group known and the machines under no leaf known; NVLink domains, with their GPUs and the trays not complete; and Sweeps.
  • Findings: each finding's severity, summary and fix, its evidence (folded) and the machines it is about.
$ astra astraeus topology
Topology of 8 machine(s), merged 2026-10-05T09:12:41.208734512Z

Fabrics (machines that reach each other over RDMA)
  FABRIC                 KIND        MACHINES
  sm-0x0002c90300a1b2c3  InfiniBand  gpu-[01-08]  found by sminfo

Tiers (spine > leaf > machines)
  spine spine-0-1  (fabric sm-0x0002c90300a1b2c3)
    leaf leaf-r0-7-g0  gpu-[01-04]  (ibnetdiscover)
    leaf leaf-r0-7-g1  gpu-[05-08]  (ibnetdiscover)

Findings (2)
  ! gpu-06 mlx5_5 trained at 200 Gb/s, its peers at 400 Gb/s: every collective over rail 5 runs at that pace
      fix: Reseat or replace the cable and transceiver of gpu-06 mlx5_5 (`mlxlink -d mlx5_5 -m` shows the link's speed and errors); until then NCCL_IB_HCA can leave the port out.
  ! Miswired: gpu-02 mlx5_3 (rail 3) is cabled to leaf-r4-g0 port 5, rail 4's leaf
      fix: Move gpu-02 mlx5_3's cable from leaf-r4-g0 port 5 to a free port of leaf-r3-g0 (rail 3's leaf).

Sweeps of the InfiniBand fabrics
  sm-0x0002c90300a1b2c3  by gpu-01 6h: 18 switches, 192 cables

! is a warning, i an info finding. Machines whose agent does not report its topology yet are listed last. --json prints the whole topology as the API answers it.

$ curl -sS "$ASTRALYX_API/topology" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
    | jq -c '.groups[] | {id, tier, parent, machines}'
{"id":"leaf-r0-7-g0","tier":0,"parent":"spine-0-1","machines":["gpu-01","gpu-02","gpu-03","gpu-04"]}
{"id":"leaf-r0-7-g1","tier":0,"parent":"spine-0-1","machines":["gpu-05","gpu-06","gpu-07","gpu-08"]}
{"id":"spine-0-1","tier":1,"parent":null,"machines":["gpu-01","gpu-02","gpu-03","gpu-04","gpu-05","gpu-06","gpu-07","gpu-08"]}

The answer has fabrics, groups, switches, links (cables between switches), machines, domains, findings and sweeps; every field is in the API reference.

Your AI assistant can read it too: the tools get_topology and get_machine_topology (Connect your AI assistant).

See one machine's topology#

Open Compute → Machines and select the machine. Two sections show its topology:

  • Its place on the network: the findings about it; then, for its fabric, spine group, leaf group, rack and rails up, what discovery Found, where From, the admin's Label, and what Placement uses; its NVLink domain; its top-of-rack switch, front-end speed and the cloud's data; and, folded, how it was found: what each source answered.
  • Inside the machine: the GPU × GPU and GPU × port matrix, with each GPU's NUMA node, PCIe link and NVLinks; and its ports: rail, state, rate against what it should train at, PCIe, NUMA, nearest GPU, the switch and port it is Cabled to, and who saw that.
$ astra astraeus topology --machine gpu-02
gpu-02  fabric sm-0x0002c90300a1b2c3  spine spine-0-1  leaf leaf-r0-7-g0  rack r12
Derived: fabric=sm-0x0002c90300a1b2c3 (sminfo), leaf=leaf-r0-7-g0 (ibnetdiscover), rack=tor-r12 (lldp), rails=8 (sysfs), spine=spine-0-1 (ibnetdiscover)
Set by an admin (these win): rack=r12

GPUs and ports (NV#: NVLinks; PIX: one PCIe switch; PXB: several bridges; PHB: the host bridge; NODE: same NUMA node; SYS: across)
        GPU0  GPU1  GPU2  GPU3  GPU4  GPU5  GPU6  GPU7  mlx5_0  mlx5_1  mlx5_2  mlx5_3  mlx5_4  mlx5_5  mlx5_6  mlx5_7  NUMA  PCIe
  GPU0  X     NV18  NV18  NV18  NV18  NV18  NV18  NV18  PIX     NODE    NODE    NODE    SYS     SYS     SYS     SYS     0     Gen5 x16
  GPU1  NV18  X     NV18  NV18  NV18  NV18  NV18  NV18  NODE    PIX     NODE    NODE    SYS     SYS     SYS     SYS     0     Gen5 x16
  GPU2  NV18  NV18  X     NV18  NV18  NV18  NV18  NV18  NODE    NODE    PIX     NODE    SYS     SYS     SYS     SYS     0     Gen5 x16
  GPU3  NV18  NV18  NV18  X     NV18  NV18  NV18  NV18  NODE    NODE    NODE    PIX     SYS     SYS     SYS     SYS     0     Gen5 x16
  GPU4  NV18  NV18  NV18  NV18  X     NV18  NV18  NV18  SYS     SYS     SYS     SYS     PIX     NODE    NODE    NODE    1     Gen5 x16
  GPU5  NV18  NV18  NV18  NV18  NV18  X     NV18  NV18  SYS     SYS     SYS     SYS     NODE    PIX     NODE    NODE    1     Gen5 x16
  GPU6  NV18  NV18  NV18  NV18  NV18  NV18  X     NV18  SYS     SYS     SYS     SYS     NODE    NODE    PIX     NODE    1     Gen5 x16
  GPU7  NV18  NV18  NV18  NV18  NV18  NV18  NV18  X     SYS     SYS     SYS     SYS     NODE    NODE    NODE    PIX     1     Gen5 x16

Ports
  PORT      RAIL  STATE  RATE      SWITCH      SWITCH PORT  SEEN BY
  mlx5_0/1  0     up     400 Gb/s  leaf-r0-g0  2            ibnetdiscover
  mlx5_1/1  1     up     400 Gb/s  leaf-r1-g0  2            ibnetdiscover
  mlx5_2/1  2     up     400 Gb/s  leaf-r2-g0  2            ibnetdiscover
  mlx5_3/1  3     up     400 Gb/s  leaf-r4-g0  5            ibnetdiscover
  mlx5_4/1  4     up     400 Gb/s  leaf-r4-g0  2            ibnetdiscover
  mlx5_5/1  5     up     400 Gb/s  leaf-r5-g0  2            ibnetdiscover
  mlx5_6/1  6     up     400 Gb/s  leaf-r6-g0  2            ibnetdiscover
  mlx5_7/1  7     up     400 Gb/s  leaf-r7-g0  2            ibnetdiscover

Findings
  ! Miswired: gpu-02 mlx5_3 (rail 3) is cabled to leaf-r4-g0 port 5, rail 4's leaf
      fix: Move gpu-02 mlx5_3's cable from leaf-r4-g0 port 5 to a free port of leaf-r3-g0 (rail 3's leaf).

A port below its rate shows both (200 Gb/s (of 400)); a PCIe link below its width, both (Gen5 x8 (of Gen5 x16)). A machine on a cloud also shows the cloud's network path and fabric.

$ curl -sS "$ASTRALYX_API/machines/gpu-02/topology" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
    | jq -c '.reported.sources[]'
{"name":"sysfs","ok":true,"detail":"8 GPU(s), 8 RDMA port(s), 8 rail(s)"}
{"name":"lldp","ok":true,"detail":"1 neighbour(s)"}
{"name":"sminfo","ok":true,"detail":"8 of 8 InfiniBand port(s) know their subnet"}
Field What
reported What the agent last read: GPUs, RDMA ports, the GPU×GPU and GPU×port matrices (gpu_gpu, gpu_nic), the top-of-rack switch (tor), the cloud's data (provider), and what each source said (sources). null for an agent that does not report its topology yet.
machine Its place in the merged topology: fabric, spine and leaf groups, rack, NVLink domain, ports and where their cables go, labels (derived, admin, sources).
placement What placement takes from it: the derived labels, its rails' count and rate, its leaf group's oversubscription, its NVLink and front-end rates.
findings The findings about it.

A machine outside the workspace's pools answers 404 NODE_NOT_FOUND.

Export to Slurm#

The topology exports as Slurm's topology.conf, for a Slurm cluster on the same machines — for example machines installed in observe-only mode beside the Slurm cluster that runs them:

Form Slurm plugin What it holds
tree (the default) topology/tree A switch per leaf group with its machines (Nodes=), a switch per spine group with its leaf groups (Switches=), and a switch named after the fabric above them where a fabric has several spine groups, or leaf groups and no spine known. Machines on a fabric under no known leaf are under <fabric>-unplaced. Fabrics that do not reach each other are separate trees.
block topology/block A block per NVLink domain (nvl-<domain>), and, for machines in none, a block per leaf group.

Machines in no switch or block are listed in a closing comment. Machine names are Astralyx's: the hosts' names, unless a machine was installed with another --name. Slurm's node names must match them. Within a workspace, the export holds only the machines of its pools.

  1. Open Compute → Topology and select Export for Slurm.
  2. Choose Tree or Block.
  3. Select Download topology.conf, or copy it.
$ astra astraeus topology --slurm tree
# topology.conf for Slurm's topology/tree, written by Astralyx from the cluster's discovered topology (2026-10-05 09:12 UTC).
# Leaf switches are leaf groups: on a rail-optimised fabric, the leaves of every rail that serve the same machines.
SwitchName=leaf-r0-7-g0 Nodes=gpu-[01-04]
SwitchName=leaf-r0-7-g1 Nodes=gpu-[05-08]
SwitchName=spine-0-1 Switches=leaf-r0-7-g0,leaf-r0-7-g1

--slurm alone is --slurm tree; --slurm block writes blocks.

$ curl -sS "$ASTRALYX_API/topology/slurm?form=block" -H "Authorization: Bearer $ASTRALYX_TOKEN"
# topology.conf for Slurm's topology/block, written by Astralyx from the cluster's discovered topology (2026-10-05 09:12 UTC).
# Blocks are NVLink domains where machines are in one, else leaf groups (one hop on every rail).
BlockName=leaf-r0-7-g0 Nodes=gpu-[01-04]
BlockName=leaf-r0-7-g1 Nodes=gpu-[05-08]

The answer is text/plain. Another form answers 400 INVALID_FORM.

Save it as Slurm's topology.conf and name the plugin in slurm.conf (TopologyPlugin=topology/tree, or topology/block).

Agent settings#

Discovery needs no setting. Three tell the machine's agent more:

Flag Variable Default Description
--topology-sweep ASTRAEUS_TOPOLOGY_SWEEP false Let this machine map its InfiniBand fabric with ibnetdiscover when asked: one machine per subnet, about once a day. See Sweeps.
--cloud <name> ASTRAEUS_CLOUD What the firmware says The cloud this machine is on, when its firmware does not say: aws, gcp, azure, oci, nebius or lambda. Its placement data is read from that cloud's metadata service (Nebius and Lambda: from --topology-file).
--topology-file <path> ASTRAEUS_TOPOLOGY_FILE none The cloud's placement data for this machine, as a file an integration keeps current, read instead of the metadata service, again every hour. See The topology file.

The installer does not write them. Set them in a systemd drop-in, which survives reruns of the installer, then restart the agent; its workers keep running:

allow-sweep.sh
sudo mkdir -p /etc/systemd/system/astraeus-agent.service.d
sudo tee /etc/systemd/system/astraeus-agent.service.d/topology.conf >/dev/null <<'EOF'
[Service]
Environment=ASTRAEUS_TOPOLOGY_SWEEP=true
EOF
sudo systemctl daemon-reload
sudo systemctl restart astraeus-agent

On AWS

The instance's role needs ec2:DescribeInstanceTopology, and the machine must reach ec2.<region>.amazonaws.com on 443 directly: that call does not go through HTTPS_PROXY. Without them, the machine's provider source says why, and the machine gets no leaf or spine from AWS.

The topology file#

--topology-file names a JSON file with the cloud's placement data for the machine, in this shape:

Field Type Description
provider string The cloud (nebius, lambda, any name). Empty: --cloud, or the cloud the firmware says, else operator. Derived names start with it: <provider>-<value>.
instance string The instance's id.
path list of strings The network above the machine, from the top down to its own switch or block. The last is its leaf (topology.astraeus.io/leaf), the one before it its spine.
fabric string Where RDMA reaches: machines with the same fabric reach each other (topology.astraeus.io/fabric, when the InfiniBand subnets do not say).
rack string The rack (topology.astraeus.io/rack, when LLDP does not say).
host string The physical host.

On Nebius, write the instance's InfiniBand topology path as path and its GPU cluster as fabric; on Lambda, the cluster as fabric, the rack as rack, and the rail group as path. Below, the values stand for your cloud's:

/etc/astraeus/topology.json
{
  "provider": "nebius",
  "instance": "<instance id>",
  "path": ["<fabric>", "<spine group>", "<leaf>"],
  "fabric": "<GPU cluster id>"
}

Name it in the agent's drop-in (Environment=ASTRAEUS_TOPOLOGY_FILE=/etc/astraeus/topology.json, as in Agent settings). The agent reads it again every hour; a file it cannot read or parse is the machine's provider source's error.

Reference: the API#

Method and path What Role
GET /topology The topology of the machines the workspace may use viewer or above
GET /topology/slurm?form=tree\|block Slurm's topology.conf, as text viewer or above
GET /machines/{name}/topology One machine's topology viewer or above
GET /runs/{name} The run, with network: where a gang went and the bandwidth to expect viewer or above

Every field is in the API reference. Fabrics are mapped again on their own (see Sweeps): a workspace cannot ask for a sweep.

Limits#

Limit Value
GPUs a machine's topology may list 64
RDMA ports a machine's topology may list 256
A sweep's duration 120 s, then stopped
Sweeps of one subnet One machine at a time; at least 30 minutes apart; again after 24 hours unchanged
LLDP read again Every 10 minutes
The cloud's data read again Every hour
A metadata request 2 s

A machine's report over these limits is refused.

Troubleshooting#

Each machine's reported.sources (GET /machines/{name}/topology) says what each source found, or why it could not.

Symptom Cause Fix
A machine is listed as not reporting its topology yet Its agent is older than topology discovery. Update its agent (Update the agent).
InfiniBand machines have a fabric but no leaf or spine No sweep: no machine of the subnet may sweep, or it cannot. Allow one or two machines per subnet (--topology-sweep) with infiniband-diags installed.
ibnetdiscover source: not installed (infiniband-diags) The machine may sweep, but the package is missing. Install infiniband-diags on that machine.
No ibnetdiscover source on the machine you allowed The setting did not reach the agent, or no InfiniBand port is up there. Check the drop-in, sudo systemctl daemon-reload, restart the agent; check ibstat.
The sweep says ibnetdiscover did not finish within 120s A very large fabric, or one that does not answer. Check the fabric with ibnetdiscover by hand on that machine.
ibnetdiscover source: no port the fabric answers (virtual functions, or no subnet manager) Virtual functions (Azure, Nebius, Google Cloud), or the subnet manager does not answer. On those clouds, the cloud's data stands in; elsewhere, check the subnet manager.
sminfo source: the subnet manager did not answer (sminfo) infiniband-diags is missing, or the subnet manager does not answer. Install infiniband-diags; check sminfo on the machine. The question is asked again after 10 minutes.
lldp source: lldpctl is not installed (lldpd) No lldpd. Install and start lldpd; the switches must send LLDP. Without it, set the rack in Where it is.
provider source: DescribeInstanceTopology: UnauthorizedOperation (the instance role needs ec2:DescribeInstanceTopology) The instance's role lacks the permission. Add ec2:DescribeInstanceTopology to the role.
provider source: the instance has no role: DescribeInstanceTopology needs one The instance has no role. Attach a role with ec2:DescribeInstanceTopology.
provider source: DescribeInstanceTopology does not list i-… AWS gives no topology for this instance type. Nothing to do: the machine gets no leaf, spine or fabric from AWS.
provider source: nebius's topology is not readable from the machine: leave it in a file (--topology-file) Nebius and Lambda are read from a file. Keep a topology file on the machine.
A rack label differs from the derived one An admin's label wins; the cabling says otherwise. Fix the label in Where it is, or empty it to use the derived rack.
An all-reduce is well below the run's expected bandwidth A degraded or miswired link on one of its machines, or the image lacks the RDMA libraries. Read the findings about its machines; check NCCL_DEBUG=INFO (Multi-machine runs).