Skip to content

Searching logs#

A run's workers write their logs on the machines they run on, and the logs stay there. When you search them, Astraeus asks each machine that has them, over the connection the machine keeps to the cluster: the machine reads the log, masks it (log masking), and sends back only the lines that match — to your request alone. Nothing of it is stored by Astraeus. Organisation admins can read a machine's own logs the same way: the kernel's (Xids, GPU resets, I/O errors) and journald's.

Search a run's logs#

The search covers every worker of the run: each one's container on the machine it runs on now and, for a worker restarted elsewhere, the log its previous attempt left on that machine.

On a run's page, Logs tab, Search every worker's log: type what to find (NCCL WARN, Traceback, CUDA error) and select Search. Each line says the worker, the machine and the line number; a machine that could not answer says why.

$ astra astraeus search-logs llm "nccl warn"
llm-1 [gpu-1 line 4211] gpu-1:1234:1301 [5] NCCL WARN NET/IB : Got async event : port error
! llm-0 on gpu-0: machine gpu-0 is not reachable right now: …
Option Default Description
<query> empty (every line) Lines containing this, in any case.
--worker <w> every worker One worker.
--tail <n> 10 000 How many of each worker's last lines are searched (at most 50 000).
--case-sensitive off Case matters.
--json The answer as JSON.
$ curl -sS "$ASTRALYX_API/runs/llm/logs?q=nccl%20warn&tail=20000" -H "Authorization: Bearer $ASTRALYX_TOKEN"
Parameter Default Description
q empty Lines containing this (at most 200 characters).
case false Case matters.
worker every worker One worker.
tail 10 000 Lines searched per worker, 1 to 50 000.

The answer is items (at most 1 000 lines: worker, machine, attempt — now or previous —, line, text), sources (what each machine read: source — container or archive —, scanned, matches, or error) and truncated when more matched.

A worker's container that is gone from its machine leaves its last 1 000 lines there, which the search reads (source: archive). A machine removed from the cluster takes its logs with it.

Read a machine's own logs#

For your organisation's admins: a machine's kernel log (dmesg) or its journal (journalctl), the last lines, those containing what you ask. Each read is audited.

On a machine's page, Machine logs: choose Kernel (dmesg) or journald, type what to find (Xid, nvme, I/O error), Read.

$ astra astraeus machine-logs dgx-1 --grep xid
2026-10-07T14:01:22,513082+00:00 NVRM: Xid (PCI:0000:1b:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
$ astra astraeus machine-logs dgx-1 --journal --unit astraeus-agent --since 2h --lines 200
$ curl -sS "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/dgx-1/system-logs?source=kernel&q=xid" -H "Authorization: Bearer $ASTRALYX_TOKEN"
Parameter Default Description
source kernel kernel or journal.
q empty Lines containing this.
lines 2 000 The last lines searched (at most 50 000).
since_minutes — The journal's last this many minutes.
unit — One systemd unit (the journal).

A machine without journald, or not reachable now, answers 503 MACHINE_LOGS_UNAVAILABLE with why. Common credential shapes are masked on the machine before the lines leave it.