Fit a big model with quantization#
You serve Llama 3.3 70B Instruct on a machine with two 48 GB GPUs (two L40S, or two RTX 6000 Ada). In BF16 the model alone is 141 GB; quantized, it fits. You work out what fits, choose the variant, deploy it across both GPUs, and trade context and parallel requests for memory.
What you need:
- A machine with two 48 GB GPUs and 60 GB of free disk at its data location.
- The variables in How the recipes are written.
1. The arithmetic#
Memory needed = weights + KV cache + overhead (Fit and quantization). For Llama 3.3 70B (80 layers, 8 KV heads of 128):
- KV cache: 2 × 80 × 8 × 128 × 2 bytes = 320 KiB per token. At 8,192 tokens × 4 requests at once, 10.7 GB; × 1 request, 2.7 GB.
- Overhead: 0.5 GB.
- Per GPU: across 2 GPUs, llama.cpp needs 5% more; each GPU's usable share is 90%.
| Variant | Weights | Needed (8K, 4 at once) | Each of 2 GPUs must have | On 2 × 48 GB |
|---|---|---|---|---|
llama3.3:70b-q4_k_m |
42.5 GB | 53.8 GB | 31.4 GB | Fits, 4 at once |
llama3.3:70b-q5_k_m |
50.0 GB | 61.2 GB | 35.7 GB | Fits, 4 at once |
llama3.3:70b-q8_0 |
75.0 GB | 86.2 GB | 50.3 GB | Only 1 at a time (78.2 GB needed, 45.6 GB per GPU) |
llama3.3:70b-bf16 (vLLM) |
141.1 GB | 145.4 GB | 80.8 GB | No: needs 4 such GPUs |
The library does this for you: open Llama 3.3, pick the machine under Runs on, and each cell says needs 2 GPUs or doesn't fit, with · 1 at a time where only fewer requests fit.
Q5_K_M is the best quality that serves four people at once here; this
recipe uses it.
2. Deploy across both GPUs#
- Eos → Models → Library → Llama 3.3, click the
Q5_K_Mcell, press Deploy. - Name:
llama70. GPUs per replica: Automatic (it chooses 2) or2. - Clear Scale to zero when idle: loading 50 GB takes a while, keep a replica up.
- Context length:
8192. Advanced → Requests at once:4. - Press Deploy.
$ curl -fsS -X POST "$CONSOLE/clusters/main/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' -d '{"catalog": "llama3.3:70b-q5_k_m"}' > /dev/null
$ curl -fsS -X POST "$API/deployments" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' \
-d '{"metadata": {"name": "llama70"}, "spec": {"model": "llama3-3-70b-q5-k-m", "gpus": 2, "context_length": 8192, "parallel": 4}}' \
> /dev/null
$ curl -fsS "$API/deployments/llama70" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
| jq '.plan | {gpus, min_gpu_memory_gb, memory_basis, context_length, parallel}'
{
"gpus": 2,
"min_gpu_memory_gb": 34,
"memory_basis": "arch",
"context_length": 8192,
"parallel": 4
}
The replica asks for 2 GPUs of one machine, each with at least the plan's
min_gpu_memory_gb (here about 34 GiB, the 35.7 GB above). llama.cpp runs
with --n-gpu-layers all --split-mode layer: half the layers on each GPU.
The first start downloads 50 GB; follow the fill run's log. Then the
deployment is Ready.
3. Trade context for memory#
The two GPUs have about 22 GB of usable memory left. Spend it on a longer context or a quantized cache:
| Change | Cache | Each GPU needs |
|---|---|---|
| As deployed: 8K × 4 | 10.7 GB | 35.7 GB |
| 16K × 4 | 21.5 GB | 42.0 GB |
| 32K × 2 | 21.5 GB | 42.0 GB |
| 32K × 4 | 42.9 GB | 54.5 GB: does not fit |
To change, Change on the deployment's page, or PUT the specification
with a new context_length and parallel. If it does not fit, the cluster
accepts the change but the new replica waits, saying it needs GPUs with more
memory; set it back.
A quantized KV cache halves the cache at a small cost in quality:
Eos's estimate still counts an f16 cache, so the replica asks for the same
GPUs; what you gain is real headroom on them, for instance to run 32K × 2
with room to spare. With vLLM, the equivalent is
["--kv-cache-dtype", "fp8"].
4. Call it#
$ curl -sS $EOS_URL/chat/completions -H "Authorization: Bearer $ASTRAEUS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "llama70", "messages": [{"role": "user", "content": "Summarise the trade-offs of 4-bit quantization."}]}'
Variations#
One 80 GB GPU. Q4_K_M (53.8 GB) or Q5_K_M (61.2 GB) fit one H100 or
A100 80 GB at 8K × 4; Automatic chooses 1 GPU.
A mixture of experts too big for the GPU. For gpt-oss-120b on a GPU
that cannot hold every expert, llama.cpp can keep some experts' weights in
the machine's RAM: engine_args: ["--n-cpu-moe", "12"] (the first 12
layers' experts on the CPU). The GPU need is then lower than Eos's
estimate; set GPUs per replica yourself.
Full precision with vLLM. llama3.3:70b-bf16 needs 4 GPUs of 48 GB (or
2 of 80 GB): "gpus": 4. vLLM splits the tensors across them.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
Pending: Waiting for a machine with 2 GPUs with at least 34 GB each, on one machine |
The two GPUs are not both free, or the machine has one. | Why? shows what holds them. A replica's GPUs are always on one machine. |
| The replica keeps exiting with an out-of-memory error in its log | Engine arguments that use more memory than the estimate (a larger --batch-size), or other processes on the GPUs. |
Remove the extra arguments, or lower context_length or parallel. |
Starting for a long time |
50 GB to download on the first start. | The fill run's log shows the bytes so far; pull the model ahead of time next time. |