Skip to content

Observe-only mode#

A machine in observe-only mode runs the Astralyx agent and is seen like any other — its GPUs and their health, its network and where it is cabled, its metrics — but Astralyx places no work there. Install it beside the scheduler that runs the machine today, Slurm or Kubernetes, to read your cluster as Astralyx sees it: the inventory, the topology and its findings, GPU faults, and the Cluster Report. When you are ready, run acceptance on some machines in a maintenance window, convert them to normal, and move their work over, one partition at a time.

What it does, and does not#

Astralyx On an observe-only machine
Reads the machine Yes, as on any machine: its GPUs, CPUs, memory, disks and network ports; GPU health and faults; its topology (PCIe, NVLink, rails, the switches it is cabled to); its metrics.
Shows the machine In Compute → Machines with the badge Observe-only, on Compute → GPUs and Compute → Topology, and in the Cluster Report with the state Observe-only.
Sees the other scheduler's work As GPU use: a GPU whose memory a Slurm job or a pod holds is Used outside Astraeus, with its utilisation, memory and temperature (GPUs).
Places work Never: no run of any workspace, no model replica, no environment or notebook, no agent run.
Checks the machine Only when an organisation admin starts checks explicitly. Not even the quick check is made by itself.
Offers acceptance for it No: observe-only machines are left out of the report's offer.

The agent only reads the machine: no container runs in its runtime until work is placed there, and none is but the checks an admin starts. That runtime, by default the containerd that comes with the agent, has its own socket and state (under /run/astraeus and /var/lib/astraeus), apart from Kubernetes' and Docker's. An observe-only machine is a machine of the cluster like any other: it counts towards Astraeus Cloud's machine limit.

Before you begin#

  • You are an organisation owner or admin: you make the join tokens and convert machines.
  • The machines meet the requirements: Linux with systemd, root access, and outbound HTTPS to Astralyx. They need no inbound port.
  • A pool of their own. A workspace sees the machines of its pools. Give the observe-only machines a pool of their own in their join token (for example pool=slurm-a, one per partition) and grant it to the workspace of the people who evaluate them (Quotas, pools and terms). Organisation admins also see every machine on the cluster's page.
  • No data location, unless you want one. A machine keeps drive data only in a data location a person chose; pass --no-data-dir and nothing is written to its disks. Acceptance then does not check its data disk: choose a data location when you convert it.

Install a machine in observe-only mode#

Add --observe-only to the install command the console gives you (see Add a machine), before &&:

$ echo '9c4e…' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --control-plane https://connect.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme Research' --org-id 0192f0c4-7d1e-7a51-9c33-5e8b2a4f6d10 --name gpu-21 --no-data-dir --observe-only && rm -f ./astraeus-token
  • The installer writes ASTRAEUS_OBSERVE_ONLY=true to /etc/astraeus/agent.env, and keeps it on later runs (an upgrade, a rerun) until you pass --no-observe-only.
  • For a fleet, add --observe-only and --no-data-dir to the installer's line in the playbook of Add a machine.
  • The machine appears in Compute → Machines, Up, with the badge Observe-only.

Note

An organisation admin's choice in Astralyx — converting the machine or making it observe-only — wins over the installer's flag, and lasts: once an admin has chosen, rerunning the installer with or without --observe-only does not change the machine's mode.

What you see#

  • Machines and GPUs. Every machine with its hardware, state and health; each GPU's health, faults, and use by the other scheduler's jobs.
  • Topology. Fabrics, leaf and spine groups, NVLink domains, and the findings with their fixes: a cable on another rail's leaf, a port below its rate, a PCIe link below its width (Topology). Export it as Slurm's topology.conf for the cluster you run today (Export to Slurm).
  • The Cluster Report. Each observe-only machine with the state Observe-only, counted in observe_only, and the results of the checks an admin ran on it.
  • Events. The machine's own: joined, down, back, conditions, and its checks' results.

Open Compute → Machines: observe-only machines carry the badge Observe-only. Select one for its page.

$ astra astraeus report --machine gpu-21

astra astraeus machines, gpus and topology list observe-only machines with the others.

$ curl -sS "$ASTRALYX_API/machines/gpu-21/validation" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>" \
  | jq '{name, observe_only, observe_only_installed, class, state: .report.state}'
{
  "name": "gpu-21",
  "observe_only": true,
  "observe_only_installed": true,
  "class": "datacenter",
  "state": "Observe-only"
}

observe_only is the mode in effect; observe_only_installed, what the installer said. In the API examples, ASTRALYX_API is https://api.astralyx.cloud/v1 and ASTRALYX_TOKEN an API token of an organisation admin for all their workspaces.

Check observe-only machines#

Checks run on an observe-only machine only when an organisation admin starts them with explicit: true (Check observe-only machines too in the console): the quick check, with DCGM's shortest diagnostic, or acceptance, the network checks or the deep suite. Without it, a request that names the machine answers 400 OBSERVE_ONLY, and one for a pool or the whole cluster leaves it out.

Drain the machines first

A check never takes a GPU another program holds: it waits for it. But the burn-in loads the machine for minutes, and the fabric checks load its network. Before acceptance, drain the machines in the scheduler that runs them, in a maintenance window:

$ scontrol update NodeName=gpu-[21-24] State=DRAIN Reason="Astralyx acceptance"

or, on Kubernetes, kubectl cordon <node> then kubectl drain <node> --ignore-daemonsets --delete-emptydir-data.

  1. Open Compute → Validation and select Run checks.
  2. Choose the suite and the machines, and turn on Check observe-only machines too.
  3. Select Start.
$ astra astraeus validate --suite acceptance --machines gpu-21,gpu-22,gpu-23,gpu-24 --explicit
validation val-20261005-9b27d4: acceptance on gpu-21, gpu-22, gpu-23, gpu-24
https://console.astralyx.cloud/o/acme/w/vision/validation?cluster=lab-a&validation=val-20261005-9b27d4
val-20261005-9b27d4

--wait follows it until it is over (CLI reference).

$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"suite": "acceptance", "machines": ["gpu-21", "gpu-22", "gpu-23", "gpu-24"], "explicit": true, "note": "slurm-a, drained"}' \
  | jq -c '{id, state, machines}'
{"id":"val-20261005-9b27d4","state":"Pending","machines":["gpu-21","gpu-22","gpu-23","gpu-24"]}

Below, <org> and <cluster> stand for your organisation's and cluster's short names.

Everything else about checks — what they measure, how they give way to work, how to read them — is in Acceptance and checks.

Move machines over#

The path from a cluster another scheduler runs to Astralyx, one partition or node pool at a time:

  1. Observe. Install the agent with --observe-only on the machines, a pool per partition. Nothing changes for the jobs running there.
  2. Read what Astralyx sees. The machines, their GPU health, the topology and its findings, the Cluster Report. Fix the miswired cables and degraded links it finds; Slurm can use the exported topology.conf right away.
  3. Accept, in a maintenance window. Drain a partition (Slurm) or a node pool (Kubernetes), then run acceptance on its machines explicitly, as above. Fix what it finds, and run it again until they pass.
  4. Convert those machines to normal (below). They take work from then on.
  5. Move the work. Submit it as runs (see Submit a run), or keep your batch scripts with astra slurm sbatch (Slurm compatibility).
  6. Keep the machines out of the other scheduler. Leave them drained (Slurm) or cordoned (Kubernetes), or take them out of its configuration: two schedulers must never place work on the same GPUs. Astralyx leaves alone a GPU another program holds, but the other scheduler does not know of Astralyx's work.

Repeat from step 3 for the next partition.

Convert a machine to normal#

Open the machine (Compute → Machines, select it) and select Convert to normal. To make a machine observe-only again, select Make observe-only.

$ astra astraeus observe-only gpu-21 off
gpu-21 is converted to normal: it takes work

astra astraeus observe-only gpu-21 on makes it observe-only again.

$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/nodes/gpu-21/observe-only" \
    -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"observe_only": false}'
{"name":"gpu-21","observe_only":false}

{"observe_only": true} makes a machine observe-only, whatever it was installed with. A body without a boolean observe_only answers 400 INVALID_REQUEST; a machine that is not your organisation's, 404 NODE_NOT_FOUND.

What converting does:

  • The machine takes work at once, like any machine of its pool.
  • Its quick check runs, within moments.
  • If its pool's policy requires acceptance, it joined after the policy was turned on, and it has not passed acceptance, it waits as Accepting until it passes.
  • The machine has the event Converted to normal by <who>: it takes work, and the change is in the organisation's audit log (cluster.node.observe_only). The machine records the choice, who made it and when (spec.scheduling.observe_only, observe_only_by as console:<e-mail>, observe_only_at).

Making a machine observe-only again places nothing new there; work already running there keeps running until it ends. Its event is Observe-only mode: set by <who> (it takes no work). Cordoning and uncordoning a machine keep its mode.

Reference#

Field Where What
spec.info.capabilities.observe_only The machine (GET /machines/{name}) true: installed with --observe-only. Absent otherwise.
spec.scheduling.observe_only The machine An organisation admin's choice, over the installer's: false converted it to normal, true made it observe-only. Absent: as installed.
spec.scheduling.observe_only_by, observe_only_at The machine Who made that choice, and when.
observe_only, observe_only_installed GET /machines/{name}/validation The mode in effect, and what the installer said.
ASTRAEUS_OBSERVE_ONLY /etc/astraeus/agent.env on the machine true with --observe-only (Installer reference).

Troubleshooting#

Symptom Cause Fix
Runs never land on the machine: machine is in observe-only mode: Astraeus watches it and places no work there The machine is observe-only. Convert it when you are ready.
400 OBSERVE_ONLY when starting checks Checks on an observe-only machine need an explicit start. Add "explicit": true (--explicit with astra, Check observe-only machines too in the console).
A step waits: gpu-21: in observe-only mode: only an admin's explicit run checks it The validation was not started explicitly. Cancel it, and start a new one with explicit: true.
An explicit check keeps waiting: its run could not be placed: the machines got busy The other scheduler's jobs hold GPUs there. Drain the machine in Slurm or Kubernetes.
The machine has no quick check Observe-only machines are not checked by themselves. Start the quick suite explicitly.
Rerunning the installer did not change the mode An organisation admin chose the mode in Astralyx; that choice wins. Convert it, or make it observe-only, in the console, with astra astraeus observe-only, or through the API.
Its GPUs show Used outside Astraeus The other scheduler's jobs use them. Nothing, while it observes. After converting, keep the machine out of the other scheduler.