Skip to content

Machine requirements#

Check this page before you add a machine to a cluster. A Linux machine needs little up front: the installer brings the container runtime and the NVIDIA container tools, and installs the few system packages that are missing. You provide the operating system, root, outbound network access and, for GPUs, a working kernel driver.

At a glance#

Requirement
Operating system Linux with systemd running and glibc 2.28 or newer (for example Ubuntu 20.04, Debian 11, RHEL 8 and later). macOS 13.3 or newer for Eos inference only.
Architecture x86-64 (x86_64, amd64) or ARM64 (aarch64, arm64). Anything else is refused.
Privileges root (through sudo). The machine agent manages containers, routes and WireGuard.
Container runtime Nothing to install: the release bundles its own containerd. Docker Engine 20.10 or newer is optional (--runtime docker).
NVIDIA GPUs The NVIDIA kernel driver. 535 or newer recommended; CUDA 12 images need 525.60.13 or newer.
AMD GPUs The kernel's amdgpu driver with /dev/kfd. ROCm is not needed on the machine.
Network Outbound HTTPS to the Astraeus address and to where the agent is downloaded from. Nothing inbound from Astraeus. See Network and firewalls.
Between machines UDP 51820 for the WireGuard mesh, when it is used.

Not supported

Distributions built on musl (Alpine), Linux without systemd (including WSL 2 without systemd enabled), 32-bit systems, and Windows. On those the installer stops with an error before it changes anything.

Operating system#

The installer checks these and stops if one is missing:

  1. You are root.
  2. systemd is running (/run/systemd/system exists). The agent runs as systemd services.
  3. The CPU architecture is x86-64 or ARM64. The installer downloads astraeus-agent-linux-amd64.tar.gz or astraeus-agent-linux-arm64.tar.gz.

The agent is built against glibc; any mainstream distribution with glibc 2.28 or newer works. The installer installs missing packages with apt, dnf, yum, zypper or pacman. With another package manager, install the packages in System packages yourself first.

Automatic NVIDIA driver installation covers Ubuntu, Debian, Fedora, RHEL, Rocky Linux, AlmaLinux and CentOS. AMD's packaged driver is offered only on Ubuntu 22.04 and 24.04 (x86-64). Elsewhere you install the driver yourself; see GPUs.

Kernel#

Feature Needed for Without it
WireGuard (wireguard module; in the mainline kernel since 5.6, backported by several distributions) The mesh: each worker gets an address of its own, and workers on different machines reach each other The machine uses the host network: two deployments on it share its ports.
overlayfs Image layers (containerd's default snapshotter) The agent falls back to the native snapshotter: slower and larger on disk.
amdgpu built with CONFIG_HSA_AMD (/dev/kfd) AMD GPUs The machine runs CPU work only.
A kernel new enough for the GPU AMD Radeon RX 7000 / Radeon PRO W7000: 6.2 or newer. Instinct MI300: 6.8 or newer. The GPU is not driven.

SELinux#

SELinux may be enforcing or permissive. On such a machine the installer labels the bundled runtime as a distribution's containerd is labelled, and installs policycoreutils-python-utils (policycoreutils-python on EL7) and container-selinux when they are missing. Without container-selinux the runtime runs as an unconfined service; the installer says so.

System packages#

The installer installs these from the distribution's repositories when they are missing:

Package Why
iproute2 (iproute on dnf/yum systems) Workers' networks and the mesh's routes. Required: the installer stops if it cannot install it.
iptables Workers' networks (NAT, published ports). Required. iptables-nft is fine.
wireguard-tools The mesh. If it cannot be installed, the machine falls back to the host network.
policycoreutils-python-utils, container-selinux Only when SELinux is enabled.

The installer does not install these; some features need them:

Package Needed for
NFS server and client (nfs-kernel-server and nfs-common on Debian/Ubuntu, nfs-utils on RHEL/Fedora) Drives that one machine serves to workers on other machines. The agent's drives part exports with exportfs and mounts NFS 4.2.
nft (nftables) A run's outbound network policy. Without it the agent writes the same rules with iptables-restore, but blocked connections are not reported.
Envoy Machines that serve external access (the ingress agent).

Container runtime#

By default the machine agent runs workers on a containerd of its own, bundled in the release and kept apart from anything else on the machine:

Component Installed in Notes
containerd 2.x (static build), ctr, the runc shim /usr/lib/astraeus/bin Runs as astraeus-containerd.service. Configuration /etc/astraeus/containerd.toml, data /var/lib/astraeus/containerd, socket /run/astraeus/containerd/containerd.sock. No TCP listener.
runc /usr/lib/astraeus/bin The OCI runtime for every worker.
gVisor (runsc, containerd-shim-runsc-v1) /usr/lib/astraeus/bin Workers with isolation: sandbox.
CNI plugins (bridge, host-local, loopback, portmap) /usr/lib/astraeus/cni Workers' network namespaces.
NVIDIA Container Toolkit (nvidia-ctk and nvidia-cdi-hook only) /usr/lib/astraeus/bin Writes the NVIDIA CDI spec. nvidia-container-runtime is not used.

None of these is put on the PATH, and nothing in /etc/containerd or /opt/containerd is read: a Docker or containerd already on the machine is left alone. Each release pins every component's version and SHA-256.

Docker Engine instead#

Pass --runtime docker to run workers on the machine's Docker Engine. Docker must already be installed (Engine API 1.41, Docker 20.10, or newer). For NVIDIA GPUs Docker also needs the NVIDIA Container Toolkit configured as a runtime (nvidia-ctk runtime configure --runtime=docker); the installer reminds you when nvidia-ctk is missing. AMD GPUs need nothing extra.

GPUs#

  • The NVIDIA kernel driver, loaded. 535 or newer is recommended; CUDA 12 images need 525.60.13 or newer.
  • Nothing else. The bundled nvidia-ctk writes the CDI spec (/var/run/cdi/nvidia.yaml) once the driver is loaded, and again when the driver version or the set of GPUs changes.
  • A GPU held by nouveau, or passed through to a VM (vfio-pci), is seen on the PCI bus but cannot be used; the machine's page says why.

The installer can install the driver on Ubuntu, Debian, Fedora, RHEL, Rocky Linux, AlmaLinux and CentOS.

  • The kernel's amdgpu driver bound to every discrete AMD GPU, and /dev/kfd present.
  • The GPU firmware (linux-firmware; firmware-amd-graphics on Debian).
  • No ROCm on the host. Container images bring it (for example rocm/pytorch, vLLM's ROCm build). No AMD container toolkit either: the agent hands each worker its GPUs' /dev/dri nodes and /dev/kfd.

Named by model: Radeon RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT, RX 7600 (XT), Radeon PRO W7900 and W7800, Instinct MI100, MI210, MI250/MI250X, MI300A, MI300X and MI325X. Other boards amdgpu drives are used too, named from the PCI database. Integrated Radeon graphics (Ryzen APUs) are ignored.

A machine without a GPU joins as a CPU machine. A machine whose GPU has no driver yet also runs CPU work, and picks the GPU up by itself once the driver is loaded.

MIG partitions and fractional GPUs are not supported: a GPU is given whole to one worker.

Disk and memory#

Astraeus enforces no minimum size. Plan for:

Path What it holds
/var/lib/astraeus/containerd Container images and their unpacked layers. Size it for the images your runs use; CUDA and ROCm images are often 10–20 GB each.
/var/lib/astraeus/worker The agent's state, each worker's files and output, the mesh key and the machine certificate.
The data location (optional, for example /mnt/nvme0/astraeus) Copies of drives kept "on each machine". Chosen at install time or later in the console.

A machine under pressure stops taking new work; what already runs keeps running.

Condition Starts when Clears when
Disk pressure Less than 10 % free on the work or image filesystem 15 % free again
Memory pressure 95 % of memory in use Below 90 %
CPU pressure (only makes the machine less preferred) 5-minute load average of 1.5 per core Below 1.2 per core

Network#

The machine only connects out. It needs:

  • HTTPS (TCP 443) to the Astraeus address (--apiserver in the install command the Add machine dialog shows).
  • HTTPS to where the agent is downloaded from: the console's /releases (--releases in the console's command), otherwise GitHub.
  • The distribution's package repositories, when packages or a driver are installed.
  • The container registries your runs pull from.
  • UDP 51820 to and from the other machines, for the mesh.

The complete list, with proxies and firewalls, is in Network and firewalls.

Time#

Keep the clock synchronised (systemd-timesyncd, chronyd or ntpd). The machine's certificate lasts 24 hours and is checked by Astraeus; a clock that is far off makes TLS fail.

macOS#

A Mac (macOS 13.3 or newer; Apple silicon for its GPU) joins as an inference machine for Eos and serves models with llama.cpp natively. It runs no Linux containers. See macOS.