Skip to content

A private coding assistant for your team#

You serve Qwen2.5 Coder as your team's coding assistant: code and questions stay on your machines. The deployment has a 32K context for whole files, serves several people at once, and adds a second replica when the team is busy. Editors and terminals connect to it with one key.

What you need:

1. Choose the size#

Eos → Models → Library, filter Category: code, open Qwen2.5 Coder. Sizes 0.5B to 32B, each in Q4_K_M (the default), Q5_K_M, Q8_0 and BF16. The library's figures at 8,192 tokens and 4 requests at once:

Variant Download Memory For
qwen2.5-coder:7b-q4_k_m 4.7 GB 7.1 GB A 24 GB GPU, with room for a long context
qwen2.5-coder:14b-q4_k_m 9.0 GB 16.0 GB A 24 GB GPU at a shorter context, or 48 GB
qwen2.5-coder:32b-q4_k_m 19.9 GB 29.0 GB A 48 GB GPU, or two 24 GB GPUs

This recipe uses the 7B. Its cache costs 56 KiB per token (28 layers, 4 KV heads of 128), so 32,768 tokens for 4 requests at once is 7.5 GB: 4.7 + 7.5 + 0.5 ≈ 12.7 GB, comfortably within a 24 GB GPU.

2. Deploy it#

  1. In Qwen2.5 Coder, click the 7B Q4_K_M cell and press Deploy.
  2. Name: coder. Clear Scale to zero when idle; Replicas, at least 1, at most 2.
  3. Context length: 32768.
  4. Advanced: Requests at once, per replica 4; Engine arguments:

    --flash-attn on
    --cache-reuse 256
    
  5. Press Deploy.

$ curl -fsS -X POST "$CONSOLE/clusters/main/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d '{"catalog": "qwen2.5-coder:7b-q4_k_m"}' > /dev/null
coder.json
{
  "metadata": {"name": "coder"},
  "spec": {
    "model": "qwen2-5-coder-7b-q4-k-m",
    "replicas": {"min": 1, "max": 2},
    "context_length": 32768,
    "parallel": 4,
    "engine_args": ["--flash-attn", "on", "--cache-reuse", "256"]
  }
}
$ curl -fsS -X POST "$API/deployments" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @coder.json > /dev/null

--cache-reuse 256 lets llama.cpp reuse the cached prefix of a prompt that changed only near its end — common when an editor resends a file. With max: 2, a second replica starts when more than 3 requests are in flight on the first (three quarters of 4), on another GPU.

The model calls tools (the library's tools): llama.cpp serves them with the model's own chat template, so agents in editors can use them.

3. Make one key for the team#

Eos → API keys → New key: Name ide-plugins, Only these: coder. Share it through your team's secret manager. Revoke and replace it when someone leaves; per-person keys work too, and show who used what in Last used.

4. Connect the editors#

In Continue's config.yaml:

~/.continue/config.yaml
name: Acme
version: 0.0.1
schema: v1
models:
  - name: Qwen2.5 Coder (Acme)
    provider: openai
    model: coder
    apiBase: https://llm.example.com/v1
    apiKey: ak-…
    roles:
      - chat
      - edit
      - apply
$ export OPENAI_API_BASE=https://llm.example.com/v1
$ export OPENAI_API_KEY=ak-…
$ aider --model openai/coder

Base URL https://llm.example.com/v1, API key ak-…, model coder.

Check from a terminal:

$ curl -sS https://llm.example.com/v1/chat/completions \
    -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
    -d '{"model": "coder", "messages": [{"role": "user", "content": "Write a Python function that parses ISO 8601 durations."}]}' \
    | jq -r '.choices[0].message.content'

5. Check that it worked#

  • The deployment's page shows Context 32,768 tokens · 4 at once and Replicas 1 serving of 1 · 1–2.
  • When several people work at once, a second replica appears under Replicas; it is removed when the load drops.
  • Organisation → Usage → Tokens by model shows the team's tokens.

Variations#

A bigger model, more GPUs. qwen2.5-coder:32b-q4_k_m on two 24 GB GPUs: set GPUs per replica to 2 (llama.cpp splits the layers across them), or leave Automatic.

vLLM for a large team. Add qwen2.5-coder:7b-bf16 and deploy it: vLLM batches up to 256 sequences per replica. For tool calls, vLLM needs the model's tool-call parser, which the library passes when it knows it; the Playground says when a deployment lacks it and offers to redeploy.

Nights and weekends. Set replicas.min to 0 and idle_minutes to 60: the GPU is free when nobody codes; the first request of the morning waits while a replica starts.

Clean up#

Delete the deployment and revoke the key. Delete the model to free its 4.7 GB on the machine.