Fine-tune a model on one GPU#
You fine-tune a language model on one GPU, as in Train on one GPU machine: Qwen2.5 1.5B Instruct with a LoRA adapter, on 4,000 instructions from the Alpaca dataset. The run merges the adapter into the base model and keeps a complete checkpoint on a drive. On a recent 24 GB GPU the training takes about fifteen minutes.
The result is ready to serve: Serve a fine-tuned model continues from it in Eos, with an OpenAI-compatible endpoint on the same machine.
flowchart LR
T[finetune-qwen<br/>run · 1 GPU] -->|adapter + merged safetensors| D1[(drive llm-runs)]
D1 -.->|next: Eos| N[Serve a fine-tuned model]
What you need:
- Everything the training recipe needs: a
workspace with one machine with a GPU (24 GB for this model), a data
location on it, the editor role, an API token,
curlandjq(How the recipes are written). - Outbound HTTPS from the machine to Hugging Face and PyPI: the run downloads the base model, the dataset and Python packages.
1. Create a drive#
A drive kept on the machine's data location, llm-runs, for everything
training produces: the cached base model and dataset, the adapter and the
merged checkpoint.
Drives → New drive, Name llm-runs, Kind On each
machine's data location, Mounted in workers at /out, Access
Read-write, Create drive.
2. Fine-tune with LoRA and merge#
The script trains a LoRA adapter, then merges it into the base model and saves a complete checkpoint (safetensors, with its tokenizer): the format both llama.cpp's converter and vLLM read. The base model and the dataset are cached on the drive, so a second run does not download them again.
import json
import os
import torch
from datasets import load_dataset
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM, AutoTokenizer
BASE = os.environ.get("BASE_MODEL", "Qwen/Qwen2.5-1.5B-Instruct")
OUT = os.environ.get("OUT_DIR", "/out")
STEPS = int(os.environ.get("STEPS", "500"))
BATCH = int(os.environ.get("BATCH_SIZE", "8"))
# ASTRAEUS_JOB_NAME is <namespace>.<run>; keep the run's own name.
RUN = os.environ.get("ASTRAEUS_JOB_NAME", "local").rsplit(".", 1)[-1]
run_dir = os.path.join(OUT, RUN)
os.makedirs(run_dir, exist_ok=True)
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, torch_dtype=torch.bfloat16).to("cuda")
model.gradient_checkpointing_enable()
model.enable_input_require_grads()
model = get_peft_model(model, LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
))
model.print_trainable_parameters()
rows = load_dataset("yahma/alpaca-cleaned", split="train[:4000]")
def as_chat(r):
prompt = r["instruction"] + (f"\n\n{r['input']}" if r["input"] else "")
return tok.apply_chat_template(
[{"role": "user", "content": prompt}, {"role": "assistant", "content": r["output"]}],
tokenize=False,
)
texts = [as_chat(r) for r in rows]
opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=2e-4)
model.train()
for step in range(STEPS):
i = (step * BATCH) % (len(texts) - BATCH)
enc = tok(texts[i:i + BATCH], return_tensors="pt", padding=True, truncation=True, max_length=1024).to("cuda")
labels = enc.input_ids.masked_fill(enc.attention_mask == 0, -100)
loss = model(**enc, labels=labels).loss
loss.backward()
opt.step()
opt.zero_grad(set_to_none=True)
if step % 25 == 0:
print(json.dumps({"step": step, "loss": round(loss.item(), 4)}), flush=True)
model.save_pretrained(os.path.join(run_dir, "adapter")) # the LoRA adapter alone
merged = model.merge_and_unload() # base + adapter, one model
merged.save_pretrained(os.path.join(run_dir, "merged"), safe_serialization=True)
tok.save_pretrained(os.path.join(run_dir, "merged"))
print(f"done: merged model in {run_dir}/merged", flush=True)
{
"metadata": {"name": "finetune-qwen", "labels": {"project": "support-bot"}},
"spec": {
"start": "Independent",
"task_template": {
"image": "pytorch/pytorch:2.4.1-cuda12.4-cudnn9-runtime",
"command": "sh",
"args": ["-c", "pip install --no-cache-dir transformers==4.46.3 peft==0.13.2 datasets==3.1.0 accelerate==1.1.1 && python /app/finetune.py"],
"env": {"OUT_DIR": "/out", "STEPS": "500", "BATCH_SIZE": "8", "HF_HOME": "/out/hf-cache"},
"restart_policy": "OnFailure",
"time_limit_seconds": 7200,
"requested_resources": {
"cpu_cores": 8,
"memory_bytes": 34359738368,
"gpu_requests": {"count": 1, "min_memory_gb": 20},
"node_selection": {"mode": "Any"}
},
"datavolume_refs": [{"name": "llm-runs", "mount_path": "/out", "mode": "ReadWrite"}],
"configs": []
}
}
}
- Put the script into the run:
jq --rawfile src finetune.py '.spec.task_template.configs = [{"mounts": ["/app/finetune.py"], "value": $src}]' finetune.json > finetune.full.json. - Runs → New run → Edit as JSON, paste
finetune.full.json, Start run.
$ jq --rawfile src finetune.py \
'.spec.task_template.configs = [{"mounts": ["/app/finetune.py"], "value": $src}]' \
finetune.json > finetune.full.json
$ curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
-H 'content-type: application/json' -d @finetune.full.json | jq -r .metadata.name
finetune-qwen
Follow it as in the training recipe:
$ astra astraeus logs finetune-qwen-0 -f
trainable params: 18,464,768 || all params: 1,562,179,072 || trainable%: 1.1820
{"step": 0, "loss": …}
{"step": 25, "loss": …}
…
done: merged model in /out/finetune-qwen/merged
The run ends Completed. On the drive llm-runs, finetune-qwen/merged/
holds config.json, model.safetensors (about 3.1 GB in BF16) and the
tokenizer; finetune-qwen/adapter/ the adapter alone (about 37 MB).
3. Check that it worked#
- The run is
Completed, and its log ends withdone: merged model in /out/finetune-qwen/merged. - The loss printed every 25 steps goes down.
- On
llm-runs,finetune-qwen/merged/holds the merged checkpoint andfinetune-qwen/adapter/the adapter.
Next: serve it#
Serve a fine-tuned model
(Eos) takes finetune-qwen/merged from here: converts it to GGUF, registers
it as a model from its drive, deploys it and calls it with the OpenAI API.
Or serve the merged folder as it is with vLLM (a variation there).
Variations#
Another base model. Any base model with a Hugging Face checkpoint works
the same way: change BASE_MODEL, and the GPU memory the run asks for.
Bigger models need a bigger GPU, or several (see
Distributed training across machines).
Your own data. Put a JSON Lines file with instruction, input and
output on the drive (or on a dataset drive) and
load it with load_dataset("json", data_files="/out/data.jsonl").
Keep only the adapter. The adapter is about 37 MB; the merged model 3.1 GB. Skip the merge to save space, but merge before serving: Eos serves complete models.
Clean up#
Delete the run, then the drive llm-runs once you have what you need.
Deleting the drive deletes the weights
Drives kept on each machine are the one kind whose files Astraeus deletes: the adapter and the merged checkpoint go with it. Copy out what you want to keep first.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
torch.OutOfMemoryError in the training log |
Batch or sequence too large for the GPU. | Lower BATCH_SIZE to 4, or max_length to 512. |
The run waits Pending |
No machine has a free GPU with 20 GB. | Open Why? next to its state; lower min_memory_gb for a smaller base model. |
| Downloads fail | No outbound HTTPS to Hugging Face or PyPI from the machine. | Allow it, or put the base model on the drive first. |