Searching logs#
A run's workers write their logs on the machines they run on, and the logs stay there. When you search them, Astraeus asks each machine that has them, over the connection the machine keeps to the cluster: the machine reads the log, masks it (log masking), and sends back only the lines that match — to your request alone. Nothing of it is stored by Astraeus. Organisation admins can read a machine's own logs the same way: the kernel's (Xids, GPU resets, I/O errors) and journald's.
Search a run's logs#
The search covers every worker of the run: each one's container on the machine it runs on now and, for a worker restarted elsewhere, the log its previous attempt left on that machine.
On a run's page, Logs tab, Search every worker's log: type what to find (NCCL WARN, Traceback, CUDA error) and select Search. Each line says the worker, the machine and the line number; a machine that could not answer says why.
$ astra astraeus search-logs llm "nccl warn"
llm-1 [gpu-1 line 4211] gpu-1:1234:1301 [5] NCCL WARN NET/IB : Got async event : port error
! llm-0 on gpu-0: machine gpu-0 is not reachable right now: …
| Option | Default | Description |
|---|---|---|
<query> |
empty (every line) | Lines containing this, in any case. |
--worker <w> |
every worker | One worker. |
--tail <n> |
10 000 | How many of each worker's last lines are searched (at most 50 000). |
--case-sensitive |
off | Case matters. |
--json |
The answer as JSON. |
$ curl -sS "$ASTRALYX_API/runs/llm/logs?q=nccl%20warn&tail=20000" -H "Authorization: Bearer $ASTRALYX_TOKEN"
| Parameter | Default | Description |
|---|---|---|
q |
empty | Lines containing this (at most 200 characters). |
case |
false |
Case matters. |
worker |
every worker | One worker. |
tail |
10 000 | Lines searched per worker, 1 to 50 000. |
The answer is items (at most 1 000 lines: worker, machine, attempt — now or previous —, line, text), sources (what each machine read: source — container or archive —, scanned, matches, or error) and truncated when more matched.
A worker's container that is gone from its machine leaves its last 1 000 lines there, which the search reads (source: archive). A machine removed from the cluster takes its logs with it.
Read a machine's own logs#
For your organisation's admins: a machine's kernel log (dmesg) or its journal (journalctl), the last lines, those containing what you ask. Each read is audited.
On a machine's page, Machine logs: choose Kernel (dmesg) or journald, type what to find (Xid, nvme, I/O error), Read.
$ curl -sS "$ASTRALYX_API/organizations/<org>/clusters/<cluster>/machines/dgx-1/system-logs?source=kernel&q=xid" -H "Authorization: Bearer $ASTRALYX_TOKEN"
| Parameter | Default | Description |
|---|---|---|
source |
kernel |
kernel or journal. |
q |
empty | Lines containing this. |
lines |
2 000 | The last lines searched (at most 50 000). |
since_minutes |
— | The journal's last this many minutes. |
unit |
— | One systemd unit (the journal). |
A machine without journald, or not reachable now, answers 503 MACHINE_LOGS_UNAVAILABLE with why. Common credential shapes are masked on the machine before the lines leave it.