A private coding assistant for your team#
You serve Qwen2.5 Coder as your team's coding assistant: code and questions stay on your machines. The deployment has a 32K context for whole files, serves several people at once, and adds a second replica when the team is busy. Editors and terminals connect to it with one key.
What you need:
- A machine with one GPU of 24 GB or more for the 7B model (or 48 GB, or two 24 GB GPUs, for the 14B), with a data location.
- A gateway your developers can reach, ideally behind TLS (An OpenAI-compatible endpoint for your app).
- The variables in How the recipes are written.
1. Choose the size#
Eos → Models → Library, filter Category: code, open Qwen2.5
Coder. Sizes 0.5B to 32B, each in Q4_K_M (the default), Q5_K_M,
Q8_0 and BF16. The library's figures at 8,192 tokens and 4 requests at
once:
| Variant | Download | Memory | For |
|---|---|---|---|
qwen2.5-coder:7b-q4_k_m |
4.7 GB | 7.1 GB | A 24 GB GPU, with room for a long context |
qwen2.5-coder:14b-q4_k_m |
9.0 GB | 16.0 GB | A 24 GB GPU at a shorter context, or 48 GB |
qwen2.5-coder:32b-q4_k_m |
19.9 GB | 29.0 GB | A 48 GB GPU, or two 24 GB GPUs |
This recipe uses the 7B. Its cache costs 56 KiB per token (28 layers, 4 KV heads of 128), so 32,768 tokens for 4 requests at once is 7.5 GB: 4.7 + 7.5 + 0.5 ≈ 12.7 GB, comfortably within a 24 GB GPU.
2. Deploy it#
- In Qwen2.5 Coder, click the 7B
Q4_K_Mcell and press Deploy. - Name:
coder. Clear Scale to zero when idle; Replicas, at least1, at most2. - Context length:
32768. -
Advanced: Requests at once, per replica
4; Engine arguments: -
Press Deploy.
$ curl -fsS -X POST "$CONSOLE/clusters/main/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' -d '{"catalog": "qwen2.5-coder:7b-q4_k_m"}' > /dev/null
--cache-reuse 256 lets llama.cpp reuse the cached prefix of a prompt that
changed only near its end — common when an editor resends a file. With
max: 2, a second replica starts when more than 3 requests are in flight
on the first (three quarters of 4), on another GPU.
The model calls tools (the library's tools): llama.cpp serves them with
the model's own chat template, so agents in editors can use them.
3. Make one key for the team#
Eos → API keys → New key: Name ide-plugins, Only these:
coder. Share it through your team's secret manager. Revoke and replace it
when someone leaves; per-person keys work too, and show who used what in
Last used.
4. Connect the editors#
In Continue's config.yaml:
Base URL https://llm.example.com/v1, API key ak-…, model coder.
Check from a terminal:
$ curl -sS https://llm.example.com/v1/chat/completions \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "coder", "messages": [{"role": "user", "content": "Write a Python function that parses ISO 8601 durations."}]}' \
| jq -r '.choices[0].message.content'
5. Check that it worked#
- The deployment's page shows Context 32,768 tokens · 4 at once and Replicas 1 serving of 1 · 1–2.
- When several people work at once, a second replica appears under Replicas; it is removed when the load drops.
- Organisation → Usage → Tokens by model shows the team's tokens.
Variations#
A bigger model, more GPUs. qwen2.5-coder:32b-q4_k_m on two 24 GB
GPUs: set GPUs per replica to 2 (llama.cpp splits the layers across
them), or leave Automatic.
vLLM for a large team. Add qwen2.5-coder:7b-bf16 and deploy it: vLLM
batches up to 256 sequences per replica. For tool calls, vLLM needs the
model's tool-call parser, which the library passes when it knows it; the
Playground says when a deployment lacks it and offers to redeploy.
Nights and weekends. Set replicas.min to 0 and idle_minutes to
60: the GPU is free when nobody codes; the first request of the morning
waits while a replica starts.
Clean up#
Delete the deployment and revoke the key. Delete the model to free its 4.7 GB on the machine.