Observe-only mode#
A machine in observe-only mode runs the Astralyx agent and is seen like any other — its GPUs and their health, its network and where it is cabled, its metrics — but Astralyx places no work there. Install it beside the scheduler that runs the machine today, Slurm or Kubernetes, to read your cluster as Astralyx sees it: the inventory, the topology and its findings, GPU faults, and the Cluster Report. When you are ready, run acceptance on some machines in a maintenance window, convert them to normal, and move their work over, one partition at a time.
What it does, and does not#
| Astralyx | On an observe-only machine |
|---|---|
| Reads the machine | Yes, as on any machine: its GPUs, CPUs, memory, disks and network ports; GPU health and faults; its topology (PCIe, NVLink, rails, the switches it is cabled to); its metrics. |
| Shows the machine | In Compute → Machines with the badge Observe-only, on Compute → GPUs and Compute → Topology, and in the Cluster Report with the state Observe-only. |
| Sees the other scheduler's work | As GPU use: a GPU whose memory a Slurm job or a pod holds is Used outside Astraeus, with its utilisation, memory and temperature (GPUs). |
| Places work | Never: no run of any workspace, no model replica, no environment or notebook, no agent run. |
| Checks the machine | Only when an organisation admin starts checks explicitly. Not even the quick check is made by itself. |
| Offers acceptance for it | No: observe-only machines are left out of the report's offer. |
The agent only reads the machine: no container runs in its runtime until
work is placed there, and none is but the checks an admin starts. That
runtime, by default the containerd that comes with the agent, has its own
socket and state (under /run/astraeus and /var/lib/astraeus), apart
from Kubernetes' and Docker's. An observe-only machine is a machine of the
cluster like any other: it counts towards Astraeus Cloud's machine limit.
Before you begin#
- You are an organisation owner or admin: you make the join tokens and convert machines.
- The machines meet the requirements: Linux with systemd, root access, and outbound HTTPS to Astralyx. They need no inbound port.
- A pool of their own. A workspace sees the machines of its pools. Give
the observe-only machines a pool of their own in their join token (for
example
pool=slurm-a, one per partition) and grant it to the workspace of the people who evaluate them (Quotas, pools and terms). Organisation admins also see every machine on the cluster's page. - No data location, unless you want one. A machine keeps drive data only
in a data location a person chose; pass
--no-data-dirand nothing is written to its disks. Acceptance then does not check its data disk: choose a data location when you convert it.
Install a machine in observe-only mode#
Add --observe-only to the install command the console gives you (see
Add a machine), before &&:
$ echo '9c4e…' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --control-plane https://connect.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme Research' --org-id 0192f0c4-7d1e-7a51-9c33-5e8b2a4f6d10 --name gpu-21 --no-data-dir --observe-only && rm -f ./astraeus-token
- The installer writes
ASTRAEUS_OBSERVE_ONLY=trueto/etc/astraeus/agent.env, and keeps it on later runs (an upgrade, a rerun) until you pass--no-observe-only. - For a fleet, add
--observe-onlyand--no-data-dirto the installer's line in the playbook of Add a machine. - The machine appears in Compute → Machines, Up, with the badge Observe-only.
Note
An organisation admin's choice in Astralyx — converting the
machine or making it observe-only —
wins over the installer's flag, and lasts: once an admin has chosen,
rerunning the installer with or without --observe-only does not change
the machine's mode.
What you see#
- Machines and GPUs. Every machine with its hardware, state and health; each GPU's health, faults, and use by the other scheduler's jobs.
- Topology. Fabrics, leaf and spine groups, NVLink domains, and the
findings with their fixes: a cable on another rail's leaf, a port below
its rate, a PCIe link below its width (Topology). Export it
as Slurm's
topology.conffor the cluster you run today (Export to Slurm). - The Cluster Report. Each observe-only machine with the state
Observe-only, counted inobserve_only, and the results of the checks an admin ran on it. - Events. The machine's own: joined, down, back, conditions, and its checks' results.
Open Compute → Machines: observe-only machines carry the badge Observe-only. Select one for its page.
astra astraeus machines, gpus and topology list observe-only
machines with the others.
$ curl -sS "$ASTRALYX_API/machines/gpu-21/validation" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Astralyx-Workspace: <org>/<workspace>" \
| jq '{name, observe_only, observe_only_installed, class, state: .report.state}'
{
"name": "gpu-21",
"observe_only": true,
"observe_only_installed": true,
"class": "datacenter",
"state": "Observe-only"
}
observe_only is the mode in effect; observe_only_installed, what
the installer said. In the API examples, ASTRALYX_API is
https://api.astralyx.cloud/v1 and ASTRALYX_TOKEN an API token of an
organisation admin for all their workspaces.
Check observe-only machines#
Checks run on an observe-only machine only when an organisation admin
starts them with explicit: true (Check observe-only machines too in
the console): the quick check, with DCGM's shortest diagnostic, or
acceptance, the network checks or the deep suite. Without it, a request
that names the machine answers 400 OBSERVE_ONLY, and one for a pool or
the whole cluster leaves it out.
Drain the machines first
A check never takes a GPU another program holds: it waits for it. But the burn-in loads the machine for minutes, and the fabric checks load its network. Before acceptance, drain the machines in the scheduler that runs them, in a maintenance window:
or, on Kubernetes, kubectl cordon <node> then
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data.
- Open Compute → Validation and select Run checks.
- Choose the suite and the machines, and turn on Check observe-only machines too.
- Select Start.
$ astra astraeus validate --suite acceptance --machines gpu-21,gpu-22,gpu-23,gpu-24 --explicit
validation val-20261005-9b27d4: acceptance on gpu-21, gpu-22, gpu-23, gpu-24
https://console.astralyx.cloud/o/acme/w/vision/validation?cluster=lab-a&validation=val-20261005-9b27d4
val-20261005-9b27d4
--wait follows it until it is over
(CLI reference).
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/validations" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"suite": "acceptance", "machines": ["gpu-21", "gpu-22", "gpu-23", "gpu-24"], "explicit": true, "note": "slurm-a, drained"}' \
| jq -c '{id, state, machines}'
{"id":"val-20261005-9b27d4","state":"Pending","machines":["gpu-21","gpu-22","gpu-23","gpu-24"]}
Below, <org> and <cluster> stand for your organisation's and
cluster's short names.
Everything else about checks — what they measure, how they give way to work, how to read them — is in Acceptance and checks.
Move machines over#
The path from a cluster another scheduler runs to Astralyx, one partition or node pool at a time:
- Observe. Install the agent with
--observe-onlyon the machines, a pool per partition. Nothing changes for the jobs running there. - Read what Astralyx sees. The machines, their GPU health, the
topology and its findings, the Cluster Report. Fix the
miswired cables and degraded links it finds; Slurm can use the exported
topology.confright away. - Accept, in a maintenance window. Drain a partition (Slurm) or a node pool (Kubernetes), then run acceptance on its machines explicitly, as above. Fix what it finds, and run it again until they pass.
- Convert those machines to normal (below). They take work from then on.
- Move the work. Submit it as runs (see Submit a run),
or keep your batch scripts with
astra slurm sbatch(Slurm compatibility). - Keep the machines out of the other scheduler. Leave them drained (Slurm) or cordoned (Kubernetes), or take them out of its configuration: two schedulers must never place work on the same GPUs. Astralyx leaves alone a GPU another program holds, but the other scheduler does not know of Astralyx's work.
Repeat from step 3 for the next partition.
Convert a machine to normal#
Open the machine (Compute → Machines, select it) and select Convert to normal. To make a machine observe-only again, select Make observe-only.
astra astraeus observe-only gpu-21 on makes it observe-only again.
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/nodes/gpu-21/observe-only" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"observe_only": false}'
{"name":"gpu-21","observe_only":false}
{"observe_only": true} makes a machine observe-only, whatever it was
installed with. A body without a boolean observe_only answers
400 INVALID_REQUEST; a machine that is not your organisation's,
404 NODE_NOT_FOUND.
What converting does:
- The machine takes work at once, like any machine of its pool.
- Its quick check runs, within moments.
- If its pool's policy requires acceptance, it joined after the policy was turned on, and it has not passed acceptance, it waits as Accepting until it passes.
- The machine has the event
Converted to normal by <who>: it takes work, and the change is in the organisation's audit log (cluster.node.observe_only). The machine records the choice, who made it and when (spec.scheduling.observe_only,observe_only_byasconsole:<e-mail>,observe_only_at).
Making a machine observe-only again places nothing new there; work already
running there keeps running until it ends. Its event is Observe-only mode:
set by <who> (it takes no work). Cordoning and uncordoning a machine keep
its mode.
Reference#
| Field | Where | What |
|---|---|---|
spec.info.capabilities.observe_only |
The machine (GET /machines/{name}) |
true: installed with --observe-only. Absent otherwise. |
spec.scheduling.observe_only |
The machine | An organisation admin's choice, over the installer's: false converted it to normal, true made it observe-only. Absent: as installed. |
spec.scheduling.observe_only_by, observe_only_at |
The machine | Who made that choice, and when. |
observe_only, observe_only_installed |
GET /machines/{name}/validation |
The mode in effect, and what the installer said. |
ASTRAEUS_OBSERVE_ONLY |
/etc/astraeus/agent.env on the machine |
true with --observe-only (Installer reference). |
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
Runs never land on the machine: machine is in observe-only mode: Astraeus watches it and places no work there |
The machine is observe-only. | Convert it when you are ready. |
400 OBSERVE_ONLY when starting checks |
Checks on an observe-only machine need an explicit start. | Add "explicit": true (--explicit with astra, Check observe-only machines too in the console). |
A step waits: gpu-21: in observe-only mode: only an admin's explicit run checks it |
The validation was not started explicitly. | Cancel it, and start a new one with explicit: true. |
An explicit check keeps waiting: its run could not be placed: the machines got busy |
The other scheduler's jobs hold GPUs there. | Drain the machine in Slurm or Kubernetes. |
| The machine has no quick check | Observe-only machines are not checked by themselves. | Start the quick suite explicitly. |
| Rerunning the installer did not change the mode | An organisation admin chose the mode in Astralyx; that choice wins. | Convert it, or make it observe-only, in the console, with astra astraeus observe-only, or through the API. |
| Its GPUs show Used outside Astraeus | The other scheduler's jobs use them. | Nothing, while it observes. After converting, keep the machine out of the other scheduler. |