Machine requirements#
Check this page before you add a machine to a cluster. A Linux machine needs little up front: the installer brings the container runtime and the NVIDIA container tools, and installs the few system packages that are missing. You provide the operating system, root, outbound network access and, for GPUs, a working kernel driver.
At a glance#
| Requirement | |
|---|---|
| Operating system | Linux with systemd running and glibc 2.28 or newer (for example Ubuntu 20.04, Debian 11, RHEL 8 and later). macOS 13.3 or newer for Eos inference only. |
| Architecture | x86-64 (x86_64, amd64) or ARM64 (aarch64, arm64). Anything else is refused. |
| Privileges | root (through sudo). The machine agent manages containers, routes and WireGuard. |
| Container runtime | Nothing to install: the release bundles its own containerd. Docker Engine 20.10 or newer is optional (--runtime docker). |
| NVIDIA GPUs | The NVIDIA kernel driver. 535 or newer recommended; CUDA 12 images need 525.60.13 or newer. |
| AMD GPUs | The kernel's amdgpu driver with /dev/kfd. ROCm is not needed on the machine. |
| Network | Outbound HTTPS to the Astraeus address and to where the agent is downloaded from. Nothing inbound from Astraeus. See Network and firewalls. |
| Between machines | UDP 51820 for the WireGuard mesh, when it is used. |
Not supported
Distributions built on musl (Alpine), Linux without systemd (including WSL 2 without systemd enabled), 32-bit systems, and Windows. On those the installer stops with an error before it changes anything.
Operating system#
The installer checks these and stops if one is missing:
- You are root.
systemdis running (/run/systemd/systemexists). The agent runs as systemd services.- The CPU architecture is x86-64 or ARM64. The installer downloads
astraeus-agent-linux-amd64.tar.gzorastraeus-agent-linux-arm64.tar.gz.
The agent is built against glibc; any mainstream distribution with glibc
2.28 or newer works. The installer installs missing packages with apt,
dnf, yum, zypper or pacman. With another package manager, install the
packages in System packages yourself first.
Automatic NVIDIA driver installation covers Ubuntu, Debian, Fedora, RHEL, Rocky Linux, AlmaLinux and CentOS. AMD's packaged driver is offered only on Ubuntu 22.04 and 24.04 (x86-64). Elsewhere you install the driver yourself; see GPUs.
Kernel#
| Feature | Needed for | Without it |
|---|---|---|
WireGuard (wireguard module; in the mainline kernel since 5.6, backported by several distributions) |
The mesh: each worker gets an address of its own, and workers on different machines reach each other | The machine uses the host network: two deployments on it share its ports. |
| overlayfs | Image layers (containerd's default snapshotter) | The agent falls back to the native snapshotter: slower and larger on disk. |
amdgpu built with CONFIG_HSA_AMD (/dev/kfd) |
AMD GPUs | The machine runs CPU work only. |
| A kernel new enough for the GPU | AMD Radeon RX 7000 / Radeon PRO W7000: 6.2 or newer. Instinct MI300: 6.8 or newer. | The GPU is not driven. |
SELinux#
SELinux may be enforcing or permissive. On such a machine the installer
labels the bundled runtime as a distribution's containerd is labelled, and
installs policycoreutils-python-utils (policycoreutils-python on EL7) and
container-selinux when they are missing. Without container-selinux the
runtime runs as an unconfined service; the installer says so.
System packages#
The installer installs these from the distribution's repositories when they are missing:
| Package | Why |
|---|---|
iproute2 (iproute on dnf/yum systems) |
Workers' networks and the mesh's routes. Required: the installer stops if it cannot install it. |
iptables |
Workers' networks (NAT, published ports). Required. iptables-nft is fine. |
wireguard-tools |
The mesh. If it cannot be installed, the machine falls back to the host network. |
policycoreutils-python-utils, container-selinux |
Only when SELinux is enabled. |
The installer does not install these; some features need them:
| Package | Needed for |
|---|---|
NFS server and client (nfs-kernel-server and nfs-common on Debian/Ubuntu, nfs-utils on RHEL/Fedora) |
Drives that one machine serves to workers on other machines. The agent's drives part exports with exportfs and mounts NFS 4.2. |
nft (nftables) |
A run's outbound network policy. Without it the agent writes the same rules with iptables-restore, but blocked connections are not reported. |
| Envoy | Machines that serve external access (the ingress agent). |
Container runtime#
By default the machine agent runs workers on a containerd of its own, bundled in the release and kept apart from anything else on the machine:
| Component | Installed in | Notes |
|---|---|---|
containerd 2.x (static build), ctr, the runc shim |
/usr/lib/astraeus/bin |
Runs as astraeus-containerd.service. Configuration /etc/astraeus/containerd.toml, data /var/lib/astraeus/containerd, socket /run/astraeus/containerd/containerd.sock. No TCP listener. |
| runc | /usr/lib/astraeus/bin |
The OCI runtime for every worker. |
gVisor (runsc, containerd-shim-runsc-v1) |
/usr/lib/astraeus/bin |
Workers with isolation: sandbox. |
CNI plugins (bridge, host-local, loopback, portmap) |
/usr/lib/astraeus/cni |
Workers' network namespaces. |
NVIDIA Container Toolkit (nvidia-ctk and nvidia-cdi-hook only) |
/usr/lib/astraeus/bin |
Writes the NVIDIA CDI spec. nvidia-container-runtime is not used. |
None of these is put on the PATH, and nothing in /etc/containerd or
/opt/containerd is read: a Docker or containerd already on the machine is
left alone. Each release pins every component's version and SHA-256.
Docker Engine instead#
Pass --runtime docker to run workers on the machine's Docker Engine. Docker
must already be installed (Engine API 1.41, Docker 20.10, or newer). For
NVIDIA GPUs Docker also needs the NVIDIA Container Toolkit configured as a
runtime (nvidia-ctk runtime configure --runtime=docker); the installer
reminds you when nvidia-ctk is missing. AMD GPUs need nothing extra.
GPUs#
- The NVIDIA kernel driver, loaded. 535 or newer is recommended; CUDA 12 images need 525.60.13 or newer.
- Nothing else. The bundled
nvidia-ctkwrites the CDI spec (/var/run/cdi/nvidia.yaml) once the driver is loaded, and again when the driver version or the set of GPUs changes. - A GPU held by
nouveau, or passed through to a VM (vfio-pci), is seen on the PCI bus but cannot be used; the machine's page says why.
The installer can install the driver on Ubuntu, Debian, Fedora, RHEL, Rocky Linux, AlmaLinux and CentOS.
- The kernel's
amdgpudriver bound to every discrete AMD GPU, and/dev/kfdpresent. - The GPU firmware (
linux-firmware;firmware-amd-graphicson Debian). - No ROCm on the host. Container images bring it (for example
rocm/pytorch, vLLM's ROCm build). No AMD container toolkit either: the agent hands each worker its GPUs'/dev/drinodes and/dev/kfd.
Named by model: Radeon RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT,
RX 7600 (XT), Radeon PRO W7900 and W7800, Instinct MI100, MI210,
MI250/MI250X, MI300A, MI300X and MI325X. Other boards amdgpu drives are
used too, named from the PCI database. Integrated Radeon graphics (Ryzen
APUs) are ignored.
A machine without a GPU joins as a CPU machine. A machine whose GPU has no driver yet also runs CPU work, and picks the GPU up by itself once the driver is loaded.
MIG partitions and fractional GPUs are not supported: a GPU is given whole to one worker.
Disk and memory#
Astraeus enforces no minimum size. Plan for:
| Path | What it holds |
|---|---|
/var/lib/astraeus/containerd |
Container images and their unpacked layers. Size it for the images your runs use; CUDA and ROCm images are often 10–20 GB each. |
/var/lib/astraeus/worker |
The agent's state, each worker's files and output, the mesh key and the machine certificate. |
The data location (optional, for example /mnt/nvme0/astraeus) |
Copies of drives kept "on each machine". Chosen at install time or later in the console. |
A machine under pressure stops taking new work; what already runs keeps running.
| Condition | Starts when | Clears when |
|---|---|---|
| Disk pressure | Less than 10 % free on the work or image filesystem | 15 % free again |
| Memory pressure | 95 % of memory in use | Below 90 % |
| CPU pressure (only makes the machine less preferred) | 5-minute load average of 1.5 per core | Below 1.2 per core |
Network#
The machine only connects out. It needs:
- HTTPS (TCP 443) to the Astraeus address (
--apiserverin the install command the Add machine dialog shows). - HTTPS to where the agent is downloaded from: the console's
/releases(--releasesin the console's command), otherwise GitHub. - The distribution's package repositories, when packages or a driver are installed.
- The container registries your runs pull from.
- UDP 51820 to and from the other machines, for the mesh.
The complete list, with proxies and firewalls, is in Network and firewalls.
Time#
Keep the clock synchronised (systemd-timesyncd, chronyd or ntpd). The
machine's certificate lasts 24 hours and is checked by Astraeus; a clock
that is far off makes TLS fail.
macOS#
A Mac (macOS 13.3 or newer; Apple silicon for its GPU) joins as an inference machine for Eos and serves models with llama.cpp natively. It runs no Linux containers. See macOS.