Skip to content

Multiple replicas behind one endpoint#

You run one deployment as several replicas on several GPUs, behind one name and one gateway address. Eos adds replicas when requests queue and removes them when the load drops; clients see one endpoint. You load-test it to watch it scale, change it without downtime, and let it scale to zero at night.

What you need:

How it works#

  • Each replica is the engine on its own GPU (or GPUs). Replicas of one deployment can be on different machines.
  • The gateway sends each request to the serving replica with the fewest requests in flight; between equals, in turn.
  • Between min and max, the deployment watches requests in flight per replica, as the engines report them. When it is above three quarters of parallel — 3 for llama.cpp's default 4 — a replica is added; when it drops, one is removed. At most one change every 2 minutes.
  • A new replica takes traffic once its health check passes.

1. Deploy with 1 to 4 replicas#

New deployment, model llama3-1-8b-q4-k-m, Name chat-fleet, clear Scale to zero when idle, Replicas, at least 1, at most 4. Press Deploy.

$ astra inference deploy llama3-1-8b-q4-k-m --name chat-fleet --min 1 --max 4
$ curl -fsS -X POST "$API/deployments" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "chat-fleet"}, "spec": {"model": "llama3-1-8b-q4-k-m", "replicas": {"min": 1, "max": 4}, "parallel": 4}}' \
    > /dev/null

The deployment starts with one replica: Replicas 1 serving of 1 · 1–4.

2. Load it#

This script keeps 16 requests in flight for five minutes:

load.py
import asyncio, os, time
from openai import AsyncOpenAI

client = AsyncOpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"], max_retries=5)
done = 0

async def worker(deadline):
    global done
    while time.time() < deadline:
        await client.chat.completions.create(
            model="chat-fleet",
            max_tokens=256,
            messages=[{"role": "user", "content": "Write a short story about a lighthouse keeper."}],
        )
        done += 1

async def main():
    deadline = time.time() + 300
    await asyncio.gather(*(worker(deadline) for _ in range(16)))
    print(f"{done} answers in 5 minutes")

asyncio.run(main())
$ python load.py

3. Watch it scale#

On the deployment's page, Replicas grows: with 16 requests in flight and 4 served at once per replica, the first replica is above 3, so a second is added; then a third and a fourth, two minutes apart, each Starting then serving. When the script ends, they are removed one by one down to 1.

$ watch -n 10 astra inference deployments
DEPLOYMENT   MODEL                KIND   STATE   REPLICAS  GATEWAY
chat-fleet   llama3-1-8b-q4-k-m   chat   Ready   3/1-4     https://llm.example.com/v1
$ curl -fsS "$API/deployments/chat-fleet" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '{replicas, ready_replicas, members: [.members[] | {machine, state, ready}]}'

If a replica cannot start because no GPU is free, it waits Pending and Why? says why; the others keep serving.

4. Change it without downtime#

Change anything — the context, the engine arguments, even the model — and replicas are replaced one at a time: a new one starts, and an old one stops only once the new one serves. With two or more replicas, clients do not notice.

$ curl -fsS "$API/deployments/chat-fleet" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq .spec > spec.json
$ jq '{spec: (. + {replicas: {min: 2, max: 4}, context_length: 16384})}' spec.json > change.json
$ curl -fsS -X PUT "$API/deployments/chat-fleet" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @change.json > /dev/null

Keep min: 2 for anything people rely on: one replica can be replaced, restarted or moved while the other serves.

5. Scale to zero at night#

For a deployment used only in working hours, let the last replica stop when idle:

"replicas": {"min": 0, "max": 4, "idle_minutes": 30}

After 30 minutes without a request the deployment is ScaledToZero and holds no GPU. The first request of the morning starts one replica and is held until it serves (up to 2 minutes at the gateway, then 503 with Retry-After; the OpenAI SDKs retry it).

Check that it worked#

  • The deployment's Replicas tab lists each replica, its machine and whether it is serving.
  • Runs lists the replicas as runs deploy-chat-fleet-0, -1… with their machines.
  • Organisation → Usage shows the GPU-hours the replicas held, and the tokens they served under the model.

Clean up#

Delete the deployment.