Serve a fine-tuned model#
You serve a model you fine-tuned on your machines with an OpenAI-compatible endpoint. The weights never leave your machines: they go from the training run's drive to a drive of their own, are converted there, and Eos serves them from it.
This recipe continues Fine-tune a model on one GPU
(Astraeus): it starts from the merged checkpoint that run left in
finetune-qwen/merged on the drive llm-runs. Any Hugging Face–format
checkpoint on a drive works the same way.
flowchart LR
D1[(drive llm-runs<br/>merged checkpoint)] --> C[convert-qwen<br/>run · CPU]
C -->|Q4_K_M GGUF| D2[(drive qwen-support)]
D2 --> M[model qwen-support<br/>from the drive]
M --> S[deployment support<br/>llama.cpp]
S --> A[Your app<br/>via the gateway]
What you need:
- The merged checkpoint from Fine-tune a model on one GPU, on the same machine.
- The editor role in the workspace, an API token,
curlandjq. - A gateway for the last step: a machine running the agent's edge part, or the hosted gateway (Gateways).
1. Create a drive for the served weights#
qwen-support holds only the weights you serve. Eos takes the size of the
served drive's copy as the size of the weights, to size the replica, so
keep nothing else in it.
Drives → New drive, Name qwen-support, Kind On each
machine's data location, Mounted in workers at /serve,
Access Read-write, Create drive.
$ curl -fsS -X POST "$API/drives" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' \
-d '{"metadata": {"name": "qwen-support"}, "spec": {"sources": [{"node_scope": "placed", "path": "", "mount_path": "/serve", "mode": "ReadWrite"}]}}' \
| jq -r .metadata.name
qwen-support
2. Convert to GGUF and quantize#
llama.cpp serves GGUF. Its own image, ghcr.io/ggml-org/llama.cpp:full,
carries the converter (convert_hf_to_gguf.py) and llama-quantize. This
run reads the merged checkpoint from llm-runs and writes only the
quantized file to qwen-support. It needs no GPU.
{
"metadata": {"name": "convert-qwen"},
"spec": {
"task_template": {
"image": "ghcr.io/ggml-org/llama.cpp:full",
"command": "sh",
"args": ["-c", "set -eu; python3 /app/convert_hf_to_gguf.py /in/finetune-qwen/merged --outtype f16 --outfile /tmp/model-f16.gguf; /app/llama-quantize /tmp/model-f16.gguf /serve/qwen-support-q4_k_m.gguf Q4_K_M; ls -l /serve"],
"restart_policy": "Never",
"requested_resources": {"cpu_cores": 4, "memory_bytes": 17179869184, "node_selection": {"mode": "Any"}},
"datavolume_refs": [
{"name": "llm-runs", "mount_path": "/in", "mode": "ReadOnly"},
{"name": "qwen-support", "mount_path": "/serve", "mode": "ReadWrite"}
]
}
}
}
$ curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' -d @convert.json > /dev/null
$ astra astraeus logs convert-qwen-0 -f | tail -n 1
-rw-r--r-- 1 root root … qwen-support-q4_k_m.gguf
The 1.5B model is about 1 GB at Q4_K_M. Use Q8_0 for more quality
(about 1.6 GB), or keep the f16 file (about 3.1 GB): name the type in the
llama-quantize call, or write the converter's output to /serve
directly.
3. Register it in Eos#
The model's weights are in your drive: register them from the drive. Nothing is downloaded or copied. The console lists such models but has no form to add one, so use the API:
{
"metadata": {"name": "qwen-support"},
"spec": {
"source": {"drive": {"name": "qwen-support", "path": "qwen-support-q4_k_m.gguf"}},
"format": "gguf",
"display": "Qwen2.5 1.5B support (LoRA, Q4_K_M)",
"params_b": 1.54,
"quantization": "Q4_K_M",
"context_length": 32768,
"capabilities": ["chat"],
"arch": {"architecture": "qwen2", "layers": 28, "kv_heads": 2, "head_dim": 128}
}
}
$ curl -fsS -X POST "$API/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' -d @model.json | jq '{name: .metadata.name, engine: .spec.engine}'
{
"name": "qwen-support",
"engine": "llama.cpp"
}
$ curl -fsS "$API/models/qwen-support" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq -r '.status | "\(.state): \(.reason)"'
Ready: Weights in drive qwen-support
| Field | Why |
|---|---|
source.drive.path |
For llama.cpp, the .gguf file itself, relative to the drive. |
format: gguf |
A drive does not say what it holds; GGUF means llama.cpp. |
context_length |
The model's maximum (Qwen2.5 1.5B: 32,768). A deployment cannot ask for more. |
arch |
The model's shape, from its config.json (num_hidden_layers, num_key_value_heads, hidden_size ÷ num_attention_heads): Eos sizes the KV cache with it. |
The model's chat template is inside the GGUF file: the converter copies the tokenizer's. Deleting this model later leaves the drive and its file alone.
4. Deploy it#
Eos → Deployments → New deployment. Model: Qwen2.5 1.5B
support (LoRA, Q4_K_M) under In this workspace. Name:
support. Clear Scale to zero when idle. Context length:
16384. Press Deploy.
The replica mounts the drive read-only at /model and runs
llama-server --model /model/qwen-support-q4_k_m.gguf … on one GPU. It is
Ready in seconds: the weights are already on the machine.
A drive kept on each machine has one copy per machine
The file is in qwen-support's copy on the machine that ran the
conversion. A replica prefers a machine that holds a copy, so on a
workspace with one GPU machine — or when that machine has room — it
lands there. If it is placed on another machine, it finds an empty copy
and its engine exits. With several machines, keep served weights on a
drive on one machine or on a shared filesystem
(Drives), or publish them to a private Hugging Face
repository (see Variations).
5. Call it#
Make an API key that may call support (Eos → API keys → New key),
then, with the base URL from the deployment's page:
$ curl -sS $EOS_URL/chat/completions \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "support", "messages": [{"role": "user", "content": "Give three tips for writing a clear bug report."}]}' \
| jq -r '.choices[0].message.content'
import os
from openai import OpenAI
client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])
reply = client.chat.completions.create(
model="support",
messages=[{"role": "user", "content": "Rewrite this politely: 'your app is broken again'"}],
)
print(reply.choices[0].message.content)
Or open Eos → Playground, pick support, and compare it with the base
model: add qwen2.5:1.5b from the library, deploy it, and use Compare.
6. Check that it worked#
- The model is
Readywith Weights in drive qwen-support; the deploymentReadywith 1 replica serving. - The replica's Log shows llama.cpp loading
/model/qwen-support-q4_k_m.gguf. - Organisation → Usage → Tokens by model shows
qwen-support.
Variations#
Serve the merged checkpoint with vLLM. Skip the conversion and register the merged folder as safetensors:
{
"metadata": {"name": "qwen-support-bf16"},
"spec": {
"source": {"drive": {"name": "qwen-support", "path": "merged"}},
"format": "safetensors",
"display": "Qwen2.5 1.5B support (BF16)",
"context_length": 32768
}
}
Write the merged checkpoint to qwen-support/merged (mount qwen-support
in the training run too and save there), so the drive holds only what is
served. vLLM reads the folder's config.json, weights and tokenizer, and
batches many requests at once.
Publish to a private Hugging Face repository. For several machines, or
to keep a versioned copy outside your cluster, push the merged checkpoint or
the GGUF file to a private repository (huggingface-cli upload
acme/qwen-support ./merged) and add it with Add from Hugging Face or a
huggingface source, naming a credential that
holds a token with read access under token. Every machine that needs it
then downloads it at the pinned commit. The weights then do leave your
machines, to Hugging Face.
Serve the adapter only. Eos serves complete models; it does not load LoRA adapters at run time. Merge first, as the fine-tuning recipe does.
Clean up#
Delete the deployment, the model (the drive stays), the conversion run, and
then the drive qwen-support.
Deleting the drive deletes the weights
Drives kept on each machine are the one kind whose files Astraeus deletes: the GGUF file goes with it. Copy out what you want to keep first.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
| The conversion fails with an unknown architecture | llama.cpp's converter does not know the base model. | Use a supported base model, or serve the merged checkpoint with vLLM. |
400 INVALID_MODEL: spec.format: gguf or safetensors (weights in a drive do not say) |
format missing. |
Add "format": "gguf". |
Creating the deployment fails with 400 INVALID_DEPLOYMENT: model qwen-support's weights are in drive qwen-support: for llama.cpp its path names the .gguf file (the first part of a split one) |
path is a folder. |
Point path at the .gguf file. |
The replica keeps exiting: failed to load model in its log |
The replica landed on a machine without the file (an empty copy), or the path is wrong. | See the warning in step 4; check the path with a run that lists /serve. |