Skip to content

Fit a big model with quantization#

You serve Llama 3.3 70B Instruct on a machine with two 48 GB GPUs (two L40S, or two RTX 6000 Ada). In BF16 the model alone is 141 GB; quantized, it fits. You work out what fits, choose the variant, deploy it across both GPUs, and trade context and parallel requests for memory.

What you need:

1. The arithmetic#

Memory needed = weights + KV cache + overhead (Fit and quantization). For Llama 3.3 70B (80 layers, 8 KV heads of 128):

  • KV cache: 2 × 80 × 8 × 128 × 2 bytes = 320 KiB per token. At 8,192 tokens × 4 requests at once, 10.7 GB; × 1 request, 2.7 GB.
  • Overhead: 0.5 GB.
  • Per GPU: across 2 GPUs, llama.cpp needs 5% more; each GPU's usable share is 90%.
Variant Weights Needed (8K, 4 at once) Each of 2 GPUs must have On 2 × 48 GB
llama3.3:70b-q4_k_m 42.5 GB 53.8 GB 31.4 GB Fits, 4 at once
llama3.3:70b-q5_k_m 50.0 GB 61.2 GB 35.7 GB Fits, 4 at once
llama3.3:70b-q8_0 75.0 GB 86.2 GB 50.3 GB Only 1 at a time (78.2 GB needed, 45.6 GB per GPU)
llama3.3:70b-bf16 (vLLM) 141.1 GB 145.4 GB 80.8 GB No: needs 4 such GPUs

The library does this for you: open Llama 3.3, pick the machine under Runs on, and each cell says needs 2 GPUs or doesn't fit, with · 1 at a time where only fewer requests fit.

Q5_K_M is the best quality that serves four people at once here; this recipe uses it.

2. Deploy across both GPUs#

  1. Eos → Models → Library → Llama 3.3, click the Q5_K_M cell, press Deploy.
  2. Name: llama70. GPUs per replica: Automatic (it chooses 2) or 2.
  3. Clear Scale to zero when idle: loading 50 GB takes a while, keep a replica up.
  4. Context length: 8192. Advanced → Requests at once: 4.
  5. Press Deploy.
$ astra inference add llama3.3:70b-q5_k_m
$ astra inference deploy llama3-3-70b-q5-k-m --name llama70 --gpus 2 --context 8192
$ curl -fsS -X POST "$CONSOLE/clusters/main/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d '{"catalog": "llama3.3:70b-q5_k_m"}' > /dev/null
$ curl -fsS -X POST "$API/deployments" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "llama70"}, "spec": {"model": "llama3-3-70b-q5-k-m", "gpus": 2, "context_length": 8192, "parallel": 4}}' \
    > /dev/null
$ curl -fsS "$API/deployments/llama70" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '.plan | {gpus, min_gpu_memory_gb, memory_basis, context_length, parallel}'
{
  "gpus": 2,
  "min_gpu_memory_gb": 34,
  "memory_basis": "arch",
  "context_length": 8192,
  "parallel": 4
}

The replica asks for 2 GPUs of one machine, each with at least the plan's min_gpu_memory_gb (here about 34 GiB, the 35.7 GB above). llama.cpp runs with --n-gpu-layers all --split-mode layer: half the layers on each GPU.

The first start downloads 50 GB; follow the fill run's log. Then the deployment is Ready.

3. Trade context for memory#

The two GPUs have about 22 GB of usable memory left. Spend it on a longer context or a quantized cache:

Change Cache Each GPU needs
As deployed: 8K × 4 10.7 GB 35.7 GB
16K × 4 21.5 GB 42.0 GB
32K × 2 21.5 GB 42.0 GB
32K × 4 42.9 GB 54.5 GB: does not fit

To change, Change on the deployment's page, or PUT the specification with a new context_length and parallel. If it does not fit, the cluster accepts the change but the new replica waits, saying it needs GPUs with more memory; set it back.

A quantized KV cache halves the cache at a small cost in quality:

"engine_args": ["--flash-attn", "on", "--cache-type-k", "q8_0", "--cache-type-v", "q8_0"]

Eos's estimate still counts an f16 cache, so the replica asks for the same GPUs; what you gain is real headroom on them, for instance to run 32K × 2 with room to spare. With vLLM, the equivalent is ["--kv-cache-dtype", "fp8"].

4. Call it#

$ curl -sS $EOS_URL/chat/completions -H "Authorization: Bearer $ASTRAEUS_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model": "llama70", "messages": [{"role": "user", "content": "Summarise the trade-offs of 4-bit quantization."}]}'

Variations#

One 80 GB GPU. Q4_K_M (53.8 GB) or Q5_K_M (61.2 GB) fit one H100 or A100 80 GB at 8K × 4; Automatic chooses 1 GPU.

A mixture of experts too big for the GPU. For gpt-oss-120b on a GPU that cannot hold every expert, llama.cpp can keep some experts' weights in the machine's RAM: engine_args: ["--n-cpu-moe", "12"] (the first 12 layers' experts on the CPU). The GPU need is then lower than Eos's estimate; set GPUs per replica yourself.

Full precision with vLLM. llama3.3:70b-bf16 needs 4 GPUs of 48 GB (or 2 of 80 GB): "gpus": 4. vLLM splits the tensors across them.

Troubleshooting#

Symptom Cause Fix
Pending: Waiting for a machine with 2 GPUs with at least 34 GB each, on one machine The two GPUs are not both free, or the machine has one. Why? shows what holds them. A replica's GPUs are always on one machine.
The replica keeps exiting with an out-of-memory error in its log Engine arguments that use more memory than the estimate (a larger --batch-size), or other processes on the GPUs. Remove the extra arguments, or lower context_length or parallel.
Starting for a long time 50 GB to download on the first start. The fill run's log shows the bytes so far; pull the model ahead of time next time.