Troubleshooting machines#
Find the symptom, check the cause, apply the fix. Most answers are in two
places: the installer's output, and the machine agent's log
(journalctl -u astraeus-agent). Start with Where the logs are
if you do not know yet what is wrong.
Where the logs are#
| What | Linux | macOS |
|---|---|---|
| Machine agent | journalctl -u astraeus-agent |
/Library/Logs/Astraeus/agent.log |
| Container runtime | journalctl -u astraeus-containerd |
— |
| Agent's drives part | journalctl -u astraeus-agent-drives |
— |
| Agent's credentials part | journalctl -u astraeus-agent-credentials |
— |
| Agent's data part | journalctl -u astraeus-agent-data |
— |
| Agent's edge part | journalctl -u astraeus-agent-edge |
— |
| A model served on a Mac | — | /var/lib/astraeus/worker/native/<id>/log |
| A worker's output | The console (the worker's page), astra astraeus logs <worker> |
Same |
| Configuration | /etc/astraeus/agent.env, /etc/astraeus/containerd.toml |
/etc/astraeus/agent.env |
Follow the agent live:
For more detail, set RUST_LOG=debug for the agent in a drop-in and restart
it; ASTRAEUS_LOG_JSON=true in agent.env switches to JSON logs:
$ sudo systemctl edit astraeus-agent
# add:
# [Service]
# Environment=RUST_LOG=debug
$ sudo systemctl restart astraeus-agent
What to collect for support#
There is no diagnostic bundle command. Collect this:
$ systemctl status astraeus-containerd astraeus-agent --no-pager
$ journalctl -u astraeus-agent -u astraeus-containerd --since "1 hour ago" --no-pager > astraeus-logs.txt
$ sudo cat /etc/astraeus/agent.env # holds the token's path, not the token
$ ip -br addr; ip route
$ sudo wg show astraeus0 # mesh machines only
$ nvidia-smi -L # NVIDIA machines
$ ls -l /var/run/cdi /etc/cdi # GPU machines
Never send /etc/astraeus/token: it is the machine's credential.
The installer stops#
The installer prints install: <reason> and exits with status 1. It changes
nothing before its checks pass.
| Message | Cause | Fix |
|---|---|---|
run as root (the worker manages containers, routes and WireGuard) |
Not run as root. | Pipe into sudo sh, as the console's command does. |
systemd is required (Linux with systemd running) |
No systemd (a container, WSL without systemd, another init). | Use a machine with systemd. |
unsupported architecture <arch> |
Not x86-64 or ARM64. | — |
--apiserver must be https:// (the token crosses the network) |
A plain http:// address. |
Use the https:// Astraeus address from the console's command. Only http://127.* and http://localhost are allowed. |
--token-file is required …, cannot read <file>, the token file is empty |
The token file is missing or empty. | Copy the whole console command, including echo '…' > ./astraeus-token. |
download failed: <url> |
The releases address is unreachable, or has no such version. | Check access to the URL (Network and firewalls); check --version; for a private GitHub repository set GITHUB_TOKEN. |
checksum mismatch: not installing |
The download does not match SHA256SUMS: a proxy altered it, or the release is incomplete. |
Retry; check what sits between the machine and the releases address. |
release <v> does not bundle containerd … |
An old release without the bundled runtime. | Install a newer version, or use --runtime docker. |
--runtime docker: Docker is not installed here … |
--runtime docker without Docker. |
Install Docker Engine, or leave out --runtime docker. |
this machine is connected already: re-run with --move … |
The machine belongs to another organisation or cluster, and there is no terminal to ask. | Pass --move to move it, or --no-move to leave it. |
nothing changed |
You answered no to the move question (or passed --no-move). |
— |
--network mesh: this kernel cannot create a WireGuard interface |
--network mesh on a kernel without WireGuard. |
Use a kernel with the wireguard module, or let the installer fall back to the host network. |
astraeus-containerd did not start (followed by its log) |
Usually SELinux labels, or overlayfs missing. | See Workers do not start. |
astraeus-agent did not start, astraeus-agent-<part> did not start |
The agent, or one of its parts, exits at once; its log follows the message. | Read the log lines printed above. |
unknown part <a> (known: drives credentials data edge), this release has no <a> part |
A wrong name in --agents. |
Use drives, credentials, data, edge (or their former names dvagent, secretsync, catalog, ingress), or none. |
The machine never appears#
The installer waits 30 seconds for the agent to say worker started. If it
does not, it prints why it is still trying:
· gpu-01 has not connected yet: … could not register the node; retrying error=…
It keeps trying: journalctl -u astraeus-agent -f
The agent keeps retrying with backoff (up to 30 seconds between tries); fix the cause and it connects by itself.
| Error in the log | Cause | Fix |
|---|---|---|
connection refused, timed out, dns error |
The Astraeus address is unreachable from the machine. | Check DNS and outbound HTTPS (TCP 443) to the Astraeus address. Check the proxy settings (Network and firewalls). |
invalid token |
The token is wrong, or was revoked. | Get a new command from Add machine. |
token expired |
The token was not used in time (24 hours for console tokens). | Get a new command. |
ENROLLMENT_USED |
Another machine used this token first. Tokens are single-use. | One token per machine. |
NODE_NAME_TAKEN |
A machine with this name is registered already, from another installation. The installer also stops with this. | Install with --name <another name>, or remove the old record (machine page → Remove machine). The machine then connects by itself. |
NODE_REMOVED |
The machine was removed from the cluster. | Run a new install command on it: it joins as a new machine. |
| A certificate or TLS error | The machine's clock is wrong, or a proxy re-signs TLS with its own CA. | Synchronise the clock. Behind a proxy that re-signs TLS, give the machine agent the proxy's CA: ASTRAEUS_APISERVER_CA=/path/proxy-ca.pem in a drop-in (sudo systemctl edit astraeus-agent). |
At token creation, Astraeus Cloud may refuse with LIMIT_REACHED: the
organisation has as many machines as it may, counting tokens not used yet.
Remove a machine, or wait for unused tokens to expire.
The machine is Down#
Down means the machine has not reported for 60 seconds. Nothing new is placed on it; its workers are Down too.
-
Is the machine on, and is the agent running?
-
Can it reach the Astraeus address? Look for
report failedorstream failedin the log: -
Did someone change the firewall or the proxy? Reports are HTTPS to the Astraeus address, every 10 seconds.
When the machine reports again it is Up at once, and workers whose containers are still running are adopted again. After an interruption on Astralyx's side, machines are given 60 seconds to report before they are marked Down.
The machine is Up but takes no work#
| Check | Where | Fix |
|---|---|---|
| Cordoned | The machine's page shows cordoned: <reason>. |
Uncordon. |
| Memory or disk pressure | Conditions on the machine's page: "Memory is nearly exhausted" (95 % used, clears below 90 %), disk under 10 % free (clears at 15 %). | Free memory or disk space (old images live in /var/lib/astraeus/containerd). |
| Not in the workspace's pools | The workspace's Settings lists its pools on the cluster. | Put the machine in a granted pool, or grant its pool (Pools). |
| Reserved | The cluster's Reservations tab. | Wait for the window, or remove the reservation. |
| GPUs fenced or busy | Compute → GPUs: Faulty, Used outside Astraeus. | See GPUs. |
| Run constraints | The run's page says why it waits. | See GPUs and placement. |
A GPU is missing#
The machine's page names the problem in Conditions and lists the GPUs on the PCI bus that are not usable.
| What the machine says | Cause | Fix |
|---|---|---|
"This machine has a GPU, but no driver for it" (DriverNotInstalled) |
No NVIDIA driver loaded. The agent log says NVIDIA GPU on the PCI bus but no NVIDIA driver loaded …: running CPU-only. |
Install the driver (535 or newer), or run the install command again: it offers to. Reboot if asked. The agent picks the GPU up by itself. |
"The GPU is held by nouveau, the open driver CUDA cannot use" (WrongDriver) |
nouveau is bound to the GPU. |
Install the NVIDIA driver (it replaces nouveau) and reboot. |
"The GPU is passed through to a virtual machine (vfio-pci)" (GPUPassedThrough) |
A VM holds it. | Release it from the VM. |
DriverUnresponsive |
An NVML call hangs: a GPU in a bad state. | Check nvidia-smi and the kernel log (dmesg, lines with Xid); reset the GPU or reboot. |
| "This machine has an AMD GPU the amdgpu driver does not drive" | amdgpu not loaded, missing from the kernel package, no firmware, or a kernel older than the GPU. |
sudo modprobe amdgpu; on Ubuntu server kernels install linux-modules-extra-$(uname -r) and linux-firmware; or run the install command again. |
AMD GPU driven, but no /dev/kfd |
The kernel was built without CONFIG_HSA_AMD. |
Use a distribution kernel, or AMD's packaged driver. |
| The GPU is listed but Faulty | A latched fault (Xid, ECC, NVLink). | Reset or replace the GPU, then Clear fault (GPUs). |
| The GPU is Used outside Astraeus | More than 10 % of its memory is used by something Astraeus did not start. | Stop that process (nvidia-smi lists it). |
Workers do not start#
A worker that cannot start shows Failed with the reason on its page.
CDI errors (GPUs)#
| Reason | Cause | Fix |
|---|---|---|
GPUs: no CDI spec names nvidia.com/gpu=GPU-…: no spec in /etc/cdi, /var/run/cdi defines nvidia.com/gpu: run nvidia-ctk cdi generate … |
The NVIDIA CDI spec was not written: the driver was not loaded when the agent last checked, or nvidia-ctk is missing from /usr/lib/astraeus/bin. |
Check nvidia-smi; restart the agent (sudo systemctl restart astraeus-agent), which writes the spec when the driver is loaded. Check /var/run/cdi/nvidia.yaml exists. |
… the specs for nvidia.com/gpu have …; the devices may have changed since the spec was written |
The GPUs changed (one fell off the bus, a driver update) since the spec was written. | Restart the agent; check nvidia-smi -L. |
CDI device … is defined by more than one spec (…): remove or regenerate the stale one |
Two files in the same CDI directory define the device (for example a hand-made /etc/cdi/nvidia.yaml and another). |
Remove the stale file. |
cdiVersion … is newer than this agent reads …: upgrade the agent |
A CDI file written by a newer tool. | Update the machine's agent. |
… failed to load: … |
A malformed spec file in /etc/cdi or /var/run/cdi. |
Fix or remove it. |
containerd#
| Symptom | Cause | Fix |
|---|---|---|
The log says containerd cannot run containers with containers and tasks services did not load |
On SELinux machines, the bundled runtime lost its labels (lib_t): containerd cannot open its store. The message names the fix when SELinux is enforcing. |
Run the installer again (it labels the runtime), or run the semanage fcontext and restorecon commands the message prints, then sudo systemctl restart astraeus-containerd. |
astraeus-containerd restarts in a loop |
Read journalctl -u astraeus-containerd. |
Common causes: a broken /etc/astraeus/containerd.toml (rerun the installer to restore it), a full /var/lib/astraeus. |
| Image pulls fail | The registry is unreachable, the image name is wrong, or credentials are missing. | Check the registry from the machine; behind a proxy, give astraeus-containerd the proxy variables (Network and firewalls). |
With --runtime docker, check systemctl status docker and that the worker
can reach /var/run/docker.sock.
Workers cannot reach each other, or names do not resolve#
| Symptom | Cause | Fix |
|---|---|---|
Runs fail with Could not resolve host on Fedora/RHEL |
firewalld drops traffic from the workers' bridges to the machine's DNS. | Run the installer again: it sets up the astraeus zone. Check firewall-cmd --get-active-zones. |
The same, with ufw or another firewall |
Incoming traffic on the astraeus bridges is dropped. |
Allow it (Other host firewalls). |
| Workers on two machines cannot reach each other | UDP 51820 is blocked between the machines, or a machine's address is not reachable from the other (different private networks). | Open UDP 51820 both ways. Check the handshake with sudo wg show astraeus0 (latest handshake). Set ASTRAEUS_NODE_IP to a reachable address (The WireGuard mesh). |
The log says mesh key unavailable; this node will not join the mesh |
wg is missing or the key cannot be written. |
Install wireguard-tools; check /var/lib/astraeus/worker/mesh. |
The log says mesh reconcile failed or node network not ready |
A wg, ip or iptables command failed; the error follows. |
Fix what the error names; the agent retries. |
The log says could not tell this machine's address; set --node-ip |
No route to the Astraeus address to take the machine's address from. | Set ASTRAEUS_NODE_IP. Until then, work on the host network cannot be reached. |
| A name resolves on one machine but not another | Machines on the host network do not use the cluster DNS. | Use the mesh, or reach workers by machine address. |
Clock and TLS#
| Log line | Meaning | Fix |
|---|---|---|
mutual TLS does not reach … using the join token |
Informational: the machine uses its token instead of a certificate, usually because a proxy ends TLS. | Nothing. To use the certificate, let the machine reach the Astraeus address without a TLS-terminating proxy. |
node certificate not usable yet; retrying |
With ASTRAEUS_NODE_IDENTITY=on, the certificate cannot be used yet. |
Check the clock, and that the Astraeus address is reached without a TLS-terminating proxy. |
| Certificate "not yet valid" or "expired" errors | The machine's clock is wrong. | Enable time synchronisation (timedatectl set-ntp true). |
An update fails#
The console or API returns UPGRADE_FAILED with the machine's reason:
| Reason | Cause | Fix |
|---|---|---|
this machine does not know where its releases are: run the installer on it once more |
The machine was installed without a releases address. | Rerun the installer. |
<url>: 404 Not Found |
The releases address has no such version (for example the sha-… version Astraeus runs, asked of GitHub releases). |
Update to a version the releases address serves, or rerun the installer with --version. |
… checksum mismatch, not installing |
The download was altered or incomplete. | Retry; check what sits between the machine and the releases address. |
| A timeout | The download took longer than 180 seconds, or the releases address is unreachable (updates do not use a proxy). | Rerun the installer on the machine instead. |
A machine still runs astraeus-worker#
Machines installed before October 2026 ran the agent as astraeus-worker
and one service per part (astraeus-dvagent, astraeus-secretsync,
astraeus-catalog, astraeus-ingress), configured by
/etc/astraeus/worker.env. An Update from the console or a rerun of the
installer moves them to astraeus-agent and astraeus-agent-<part>
(Agent names before October 2026). If
systemctl status astraeus-worker still shows it running after an update:
-
Look at what the agent said when it started under its former name:
Message What it means Fix moving this machine to the agent's services (astraeus-agent); the switch follows in a few secondsThe switch was scheduled. Wait a few seconds, then check systemctl is-active astraeus-agent.could not move this machine to the agent's services: it keeps running as it is (the installer moves it)The switch could not be prepared (the reason follows in the message). The machine keeps working under the former names. Run the installer again. Nothing The agent was not started as root under systemd, or the release installed is older than the one binary. Update to a current release, or run the installer again. -
If the switch ran but
astraeus-agentdid not start, the former services were started again andastraeus-agentwas disabled. Read why it did not start, and the switch's own log:$ journalctl -u astraeus-agent --since "1 hour ago" --no-pager $ journalctl -u astraeus-agent-move --no-pagerFix the cause, then run the installer again.
Running work keeps running throughout: stopping either service leaves the
containers running. Your own systemd drop-ins are not all carried over: an
update copies the main service's (astraeus-worker.service.d) to
astraeus-agent.service.d, a rerun of the installer does not, and neither
copies those of the parts. After the move, check them with systemctl cat
astraeus-agent and recreate the ones you need under the new names (for
example a proxy).
A Mac keeps its former names after an Update; run the installer again to move it (macOS).
macOS#
| Symptom | Fix |
|---|---|
| The Mac drops off when idle | It slept. Keep it on power and the lid open, or sudo pmset -a sleep 0. |
| The agent does not run after a restart | Turn on Allow in the Background for it in System Settings → General → Login Items. |
--data-dir: "macOS protects …" |
Use a folder outside Documents, Desktop, Downloads, iCloud and /Volumes, such as /Users/Shared/Astraeus. |
| A run is refused with "Linux containers on a Mac arrive with the VM" | Macs run Eos models only. See macOS. |