Skip to content

Serve a fine-tuned model#

You serve a model you fine-tuned on your machines with an OpenAI-compatible endpoint. The weights never leave your machines: they go from the training run's drive to a drive of their own, are converted there, and Eos serves them from it.

This recipe continues Fine-tune a model on one GPU (Astraeus): it starts from the merged checkpoint that run left in finetune-qwen/merged on the drive llm-runs. Any Hugging Face–format checkpoint on a drive works the same way.

flowchart LR
  D1[(drive llm-runs<br/>merged checkpoint)] --> C[convert-qwen<br/>run · CPU]
  C -->|Q4_K_M GGUF| D2[(drive qwen-support)]
  D2 --> M[model qwen-support<br/>from the drive]
  M --> S[deployment support<br/>llama.cpp]
  S --> A[Your app<br/>via the gateway]

What you need:

  • The merged checkpoint from Fine-tune a model on one GPU, on the same machine.
  • The editor role in the workspace, an API token, curl and jq.
  • A gateway for the last step: a machine running the agent's edge part, or the hosted gateway (Gateways).

1. Create a drive for the served weights#

qwen-support holds only the weights you serve. Eos takes the size of the served drive's copy as the size of the weights, to size the replica, so keep nothing else in it.

Drives → New drive, Name qwen-support, Kind On each machine's data location, Mounted in workers at /serve, Access Read-write, Create drive.

$ curl -fsS -X POST "$API/drives" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "qwen-support"}, "spec": {"sources": [{"node_scope": "placed", "path": "", "mount_path": "/serve", "mode": "ReadWrite"}]}}' \
    | jq -r .metadata.name
qwen-support

2. Convert to GGUF and quantize#

llama.cpp serves GGUF. Its own image, ghcr.io/ggml-org/llama.cpp:full, carries the converter (convert_hf_to_gguf.py) and llama-quantize. This run reads the merged checkpoint from llm-runs and writes only the quantized file to qwen-support. It needs no GPU.

convert.json
{
  "metadata": {"name": "convert-qwen"},
  "spec": {
    "task_template": {
      "image": "ghcr.io/ggml-org/llama.cpp:full",
      "command": "sh",
      "args": ["-c", "set -eu; python3 /app/convert_hf_to_gguf.py /in/finetune-qwen/merged --outtype f16 --outfile /tmp/model-f16.gguf; /app/llama-quantize /tmp/model-f16.gguf /serve/qwen-support-q4_k_m.gguf Q4_K_M; ls -l /serve"],
      "restart_policy": "Never",
      "requested_resources": {"cpu_cores": 4, "memory_bytes": 17179869184, "node_selection": {"mode": "Any"}},
      "datavolume_refs": [
        {"name": "llm-runs", "mount_path": "/in", "mode": "ReadOnly"},
        {"name": "qwen-support", "mount_path": "/serve", "mode": "ReadWrite"}
      ]
    }
  }
}
$ curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @convert.json > /dev/null
$ astra astraeus logs convert-qwen-0 -f | tail -n 1
-rw-r--r-- 1 root root … qwen-support-q4_k_m.gguf

The 1.5B model is about 1 GB at Q4_K_M. Use Q8_0 for more quality (about 1.6 GB), or keep the f16 file (about 3.1 GB): name the type in the llama-quantize call, or write the converter's output to /serve directly.

3. Register it in Eos#

The model's weights are in your drive: register them from the drive. Nothing is downloaded or copied. The console lists such models but has no form to add one, so use the API:

model.json
{
  "metadata": {"name": "qwen-support"},
  "spec": {
    "source": {"drive": {"name": "qwen-support", "path": "qwen-support-q4_k_m.gguf"}},
    "format": "gguf",
    "display": "Qwen2.5 1.5B support (LoRA, Q4_K_M)",
    "params_b": 1.54,
    "quantization": "Q4_K_M",
    "context_length": 32768,
    "capabilities": ["chat"],
    "arch": {"architecture": "qwen2", "layers": 28, "kv_heads": 2, "head_dim": 128}
  }
}
$ curl -fsS -X POST "$API/models" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @model.json | jq '{name: .metadata.name, engine: .spec.engine}'
{
  "name": "qwen-support",
  "engine": "llama.cpp"
}
$ curl -fsS "$API/models/qwen-support" -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq -r '.status | "\(.state): \(.reason)"'
Ready: Weights in drive qwen-support
Field Why
source.drive.path For llama.cpp, the .gguf file itself, relative to the drive.
format: gguf A drive does not say what it holds; GGUF means llama.cpp.
context_length The model's maximum (Qwen2.5 1.5B: 32,768). A deployment cannot ask for more.
arch The model's shape, from its config.json (num_hidden_layers, num_key_value_heads, hidden_size ÷ num_attention_heads): Eos sizes the KV cache with it.

The model's chat template is inside the GGUF file: the converter copies the tokenizer's. Deleting this model later leaves the drive and its file alone.

4. Deploy it#

Eos → Deployments → New deployment. Model: Qwen2.5 1.5B support (LoRA, Q4_K_M) under In this workspace. Name: support. Clear Scale to zero when idle. Context length: 16384. Press Deploy.

$ astra inference deploy qwen-support --name support --context 16384
$ curl -fsS -X POST "$API/deployments" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "support"}, "spec": {"model": "qwen-support", "replicas": {"min": 1, "max": 1}, "context_length": 16384}}' \
    > /dev/null

The replica mounts the drive read-only at /model and runs llama-server --model /model/qwen-support-q4_k_m.gguf … on one GPU. It is Ready in seconds: the weights are already on the machine.

A drive kept on each machine has one copy per machine

The file is in qwen-support's copy on the machine that ran the conversion. A replica prefers a machine that holds a copy, so on a workspace with one GPU machine — or when that machine has room — it lands there. If it is placed on another machine, it finds an empty copy and its engine exits. With several machines, keep served weights on a drive on one machine or on a shared filesystem (Drives), or publish them to a private Hugging Face repository (see Variations).

5. Call it#

Make an API key that may call support (Eos → API keys → New key), then, with the base URL from the deployment's page:

$ curl -sS $EOS_URL/chat/completions \
    -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
    -d '{"model": "support", "messages": [{"role": "user", "content": "Give three tips for writing a clear bug report."}]}' \
    | jq -r '.choices[0].message.content'
ask.py
import os
from openai import OpenAI

client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])
reply = client.chat.completions.create(
    model="support",
    messages=[{"role": "user", "content": "Rewrite this politely: 'your app is broken again'"}],
)
print(reply.choices[0].message.content)

Or open Eos → Playground, pick support, and compare it with the base model: add qwen2.5:1.5b from the library, deploy it, and use Compare.

6. Check that it worked#

  • The model is Ready with Weights in drive qwen-support; the deployment Ready with 1 replica serving.
  • The replica's Log shows llama.cpp loading /model/qwen-support-q4_k_m.gguf.
  • Organisation → Usage → Tokens by model shows qwen-support.

Variations#

Serve the merged checkpoint with vLLM. Skip the conversion and register the merged folder as safetensors:

{
  "metadata": {"name": "qwen-support-bf16"},
  "spec": {
    "source": {"drive": {"name": "qwen-support", "path": "merged"}},
    "format": "safetensors",
    "display": "Qwen2.5 1.5B support (BF16)",
    "context_length": 32768
  }
}

Write the merged checkpoint to qwen-support/merged (mount qwen-support in the training run too and save there), so the drive holds only what is served. vLLM reads the folder's config.json, weights and tokenizer, and batches many requests at once.

Publish to a private Hugging Face repository. For several machines, or to keep a versioned copy outside your cluster, push the merged checkpoint or the GGUF file to a private repository (huggingface-cli upload acme/qwen-support ./merged) and add it with Add from Hugging Face or a huggingface source, naming a credential that holds a token with read access under token. Every machine that needs it then downloads it at the pinned commit. The weights then do leave your machines, to Hugging Face.

Serve the adapter only. Eos serves complete models; it does not load LoRA adapters at run time. Merge first, as the fine-tuning recipe does.

Clean up#

Delete the deployment, the model (the drive stays), the conversion run, and then the drive qwen-support.

Deleting the drive deletes the weights

Drives kept on each machine are the one kind whose files Astraeus deletes: the GGUF file goes with it. Copy out what you want to keep first.

Troubleshooting#

Symptom Cause Fix
The conversion fails with an unknown architecture llama.cpp's converter does not know the base model. Use a supported base model, or serve the merged checkpoint with vLLM.
400 INVALID_MODEL: spec.format: gguf or safetensors (weights in a drive do not say) format missing. Add "format": "gguf".
Creating the deployment fails with 400 INVALID_DEPLOYMENT: model qwen-support's weights are in drive qwen-support: for llama.cpp its path names the .gguf file (the first part of a split one) path is a folder. Point path at the .gguf file.
The replica keeps exiting: failed to load model in its log The replica landed on a machine without the file (an empty copy), or the path is wrong. See the warning in step 4; check the path with a run that lists /serve.