Skip to content

Deployments and engines#

A deployment serves one model: replicas of an inference engine on your machines, behind one endpoint. Clients speak the OpenAI API to it and put the deployment's name in a request's model. This page explains what a replica is, which engine to choose, how GPUs and CPUs are used, and how a deployment scales.

Replicas#

Each replica is one engine process with its own copy of the model in memory, running on one machine:

  • On GPUs: 1, 2, 4 or 8 whole GPUs of one machine. A GPU is never shared between replicas. With several GPUs, llama.cpp splits the model's layers across them and vLLM splits its tensors (tensor parallelism).
  • On CPUs: for a GGUF model small enough to answer from RAM. llama.cpp only.

A replica is a run of your workspace, like any other: it waits in the queue for a machine, counts against your quota, and shows in Runs with its log. It mounts the model's weights read-only; if the machine has no copy yet, it downloads them first.

A replica takes traffic once its engine answers its health check. All serving replicas sit behind one endpoint; requests go to the replica with the fewest requests in flight.

flowchart LR
  C[Your app] -->|"Authorization: Bearer ak-…<br/>model: chat"| G[Gateway<br/>on an edge machine]
  G --> R1[Replica 1<br/>gpu-01 · 1 GPU]
  G --> R2[Replica 2<br/>gpu-02 · 1 GPU]
  W1[(Weights copy<br/>gpu-01)] -.-> R1
  W2[(Weights copy<br/>gpu-02)] -.-> R2

Choosing the engine#

A deployment's engine follows its model's format unless you choose:

llama.cpp vLLM
Serves GGUF files safetensors checkpoints (Hugging Face layout)
Runs on NVIDIA GPUs (CUDA), AMD GPUs (ROCm, or Vulkan), a Mac's GPU (Metal), CPUs NVIDIA GPUs (CUDA), AMD GPUs its ROCm build supports
Quantization GGUF quantizations: Q4_K_M, Q5_K_M, Q8_0, F16… chosen by the file What the checkpoint is: BF16, FP16, FP8, MXFP4…
Memory The KV cache for every parallel request is allocated at start Takes 90% of each GPU and pages its KV cache between requests
Requests at once (default) 4 on a GPU (2 or 1 when that is what fits), 2 on a CPU 256
Best for One GPU or a workstation, small and mid-size models, quantized models, CPUs, Macs, AMD Radeon cards Many concurrent users, high throughput, full-precision checkpoints, large GPUs
Embeddings --embeddings with the model's pooling --runner pooling --convert embed
Speech to text, video — Whisper; videos in a conversation

Choose llama.cpp when memory is tight (Llama 3.1 70B at Q4_K_M is 42.5 GB of weights, against about 140 GB in BF16), when you serve on CPUs or a Mac, or when a handful of people use the model at once.

Choose vLLM when many requests arrive together: its paged cache and continuous batching serve far more requests per GPU. It needs the full checkpoint (or a vLLM-quantized one such as FP8) and a GPU with room for it.

The library offers both for most families: GGUF variants (llama.cpp) and a bf16 variant (vLLM). See Choose the engine and quantization.

The engines are pinned: llama.cpp release v0.5.0 and vLLM v0.30.0, each image by digest, so what serves your model never changes under it. Every flag a deployment passes is listed in the Deployment specification.

GPU vendors#

A deployment runs on any GPU its engine runs on, unless it names a vendor (gpu_vendor, the console's GPU maker). Each machine runs the engine built for its own GPUs, with the same release and the same arguments:

Machine llama.cpp vLLM
NVIDIA GPUs CUDA build CUDA build
AMD Instinct MI300X/A, MI325X, MI210/MI250 ROCm build ROCm build
AMD Instinct MI100 ROCm build —
AMD Instinct MI350/MI355 — ROCm build
AMD Radeon RX 7900 XTX/XT/GRE, PRO W7900/W7800, RX 7800/7700 ROCm build ROCm build
AMD Radeon RX 9070/9060, Ryzen AI (gfx1150, gfx1151) ROCm build ROCm build
AMD Radeon RX 7600, RX 6800/6900 ROCm build —
Any other AMD Radeon (RX 6700, 6600, 5000…) Vulkan build —
Apple silicon Mac Native Metal build —
CPU only CPU build —

A replica is placed only on a machine whose GPUs its engine runs on; the others are listed with the reason. See Serve on AMD GPUs and macOS.

AMD support

The AMD builds and their GPU lists come from the pinned releases' own build files. Serving on AMD hardware is newer than serving on NVIDIA hardware; report what you see.

GPUs per replica#

Leave GPUs per replica on Automatic and Eos chooses the fewest GPUs that hold the model on the machine with the largest GPUs your workspace may use — and, for llama.cpp, the most requests at once (4, then 2, then 1) that still fit there. The choice is written into the deployment when you create or change it, so it does not move as machines come and go. To decide yourself, set 1, 2, 4 or 8. See Fit and quantization.

Context and parallel requests#

  • Context length is the longest context one request may use, in tokens. By default it is the model's, at most 8192 for llama.cpp (its cache is allocated whole at start, for every parallel slot) and at most 32768 for vLLM. You may ask for more, up to the model's own maximum.
  • Requests at once (parallel) is how many requests one replica serves together: llama.cpp's slots, vLLM's maximum batched sequences. With llama.cpp, each slot has the whole context, so memory grows with context × parallel.

A longer context or more parallel requests need more GPU memory; a shorter context or a smaller quantization need less.

Scaling#

Setting What happens
min = max A fixed number of replicas.
min < max Replicas are added when requests queue — more in flight per replica than three quarters of what one serves at once — and removed when they drain, at most once every 2 minutes.
min: 0 (Scale to zero when idle) After 15 minutes without a request (1 to 1440, your choice) the last replica stops and the deployment is ScaledToZero: it holds no GPU and costs nothing. The next request starts a replica and is held until it serves.

A request to a deployment at zero waits while a replica starts: the weights are already on the machine after the first time, but the engine still loads them, which takes from seconds to a few minutes for large models. Through the gateway on your machines the request is held up to 2 minutes, then answered 503 with Retry-After: 5. Keep min: 1 for anything interactive.

Changing a deployment#

You can change the replicas, context, parallel requests, GPUs, engine arguments and the other fields at any time. Replicas are replaced one at a time: a new one starts, and an old one stops only once the new one serves. Deleting a deployment stops its replicas and removes its endpoint; the model and its weights stay.

States#

State Meaning
Pending Waiting for a machine with enough free GPUs (or CPUs), or for the weights. The reason says what for, for example Waiting for a machine with a GPU with at least 15 GB.
Starting A replica is downloading the weights or the engine is loading them.
Ready At least one replica serves: 1 replica serving, 1 of 3 replicas serving.
ScaledToZero No replica runs; the next request starts one.
Failed It cannot serve as asked: the model is gone or its download failed, the specification does not fit the model, or a replica keeps exiting.

When a deployment or one of its replicas waits, the console's Why? explains it machine by machine. See Troubleshoot a deployment.

Reaching a deployment#

From Address Key
Your applications, anywhere A gateway: on one of your machines (http://<machine>:8800/v1, or the address you put in front of it), or the hosted gateway An API key
Runs of the same workspace http://deploy-<name>.<namespace>.astraeus.local:8000/v1, shown on the deployment's page under Endpoint None
People The Playground and Try it Their sign-in

See Gateways, API keys, sharing and usage.