Multiple replicas behind one endpoint#
You run one deployment as several replicas on several GPUs, behind one name and one gateway address. Eos adds replicas when requests queue and removes them when the load drops; clients see one endpoint. You load-test it to watch it scale, change it without downtime, and let it scale to zero at night.
What you need:
- GPUs for up to 4 replicas: four GPUs of 16 GB or more, on one machine or several.
- A model in the workspace, here
llama3-1-8b-q4-k-m(Serve an open chat model on one GPU machine). - A key, a gateway, Python with
openai, and the variables in How the recipes are written.
How it works#
- Each replica is the engine on its own GPU (or GPUs). Replicas of one deployment can be on different machines.
- The gateway sends each request to the serving replica with the fewest requests in flight; between equals, in turn.
- Between
minandmax, the deployment watches requests in flight per replica, as the engines report them. When it is above three quarters ofparallel— 3 for llama.cpp's default 4 — a replica is added; when it drops, one is removed. At most one change every 2 minutes. - A new replica takes traffic once its health check passes.
1. Deploy with 1 to 4 replicas#
New deployment, model llama3-1-8b-q4-k-m, Name chat-fleet,
clear Scale to zero when idle, Replicas, at least 1, at
most 4. Press Deploy.
The deployment starts with one replica: Replicas 1 serving of 1 · 1–4.
2. Load it#
This script keeps 16 requests in flight for five minutes:
import asyncio, os, time
from openai import AsyncOpenAI
client = AsyncOpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"], max_retries=5)
done = 0
async def worker(deadline):
global done
while time.time() < deadline:
await client.chat.completions.create(
model="chat-fleet",
max_tokens=256,
messages=[{"role": "user", "content": "Write a short story about a lighthouse keeper."}],
)
done += 1
async def main():
deadline = time.time() + 300
await asyncio.gather(*(worker(deadline) for _ in range(16)))
print(f"{done} answers in 5 minutes")
asyncio.run(main())
3. Watch it scale#
On the deployment's page, Replicas grows: with 16 requests in flight
and 4 served at once per replica, the first replica is above 3, so a
second is added; then a third and a fourth, two minutes apart, each
Starting then serving. When the script ends, they are removed one by
one down to 1.
If a replica cannot start because no GPU is free, it waits Pending and
Why? says why; the others keep serving.
4. Change it without downtime#
Change anything — the context, the engine arguments, even the model — and replicas are replaced one at a time: a new one starts, and an old one stops only once the new one serves. With two or more replicas, clients do not notice.
$ curl -fsS "$API/deployments/chat-fleet" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq .spec > spec.json
$ jq '{spec: (. + {replicas: {min: 2, max: 4}, context_length: 16384})}' spec.json > change.json
$ curl -fsS -X PUT "$API/deployments/chat-fleet" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' -d @change.json > /dev/null
Keep min: 2 for anything people rely on: one replica can be replaced,
restarted or moved while the other serves.
5. Scale to zero at night#
For a deployment used only in working hours, let the last replica stop when idle:
After 30 minutes without a request the deployment is ScaledToZero and
holds no GPU. The first request of the morning starts one replica and is
held until it serves (up to 2 minutes at the gateway, then 503 with
Retry-After; the OpenAI SDKs retry it).
Check that it worked#
- The deployment's Replicas tab lists each replica, its machine and whether it is serving.
- Runs lists the replicas as runs
deploy-chat-fleet-0,-1… with their machines. - Organisation → Usage shows the GPU-hours the replicas held, and the tokens they served under the model.
Clean up#
Delete the deployment.