Skip to content

Fine-tune a model on one GPU#

You fine-tune a language model on one GPU, as in Train on one GPU machine: Qwen2.5 1.5B Instruct with a LoRA adapter, on 4,000 instructions from the Alpaca dataset. The run merges the adapter into the base model and keeps a complete checkpoint on a drive. On a recent 24 GB GPU the training takes about fifteen minutes.

The result is ready to serve: Serve a fine-tuned model continues from it in Eos, with an OpenAI-compatible endpoint on the same machine.

flowchart LR
  T[finetune-qwen<br/>run · 1 GPU] -->|adapter + merged safetensors| D1[(drive llm-runs)]
  D1 -.->|next: Eos| N[Serve a fine-tuned model]

What you need:

  • Everything the training recipe needs: a workspace with one machine with a GPU (24 GB for this model), a data location on it, the editor role, an API token, curl and jq (How the recipes are written).
  • Outbound HTTPS from the machine to Hugging Face and PyPI: the run downloads the base model, the dataset and Python packages.

1. Create a drive#

A drive kept on the machine's data location, llm-runs, for everything training produces: the cached base model and dataset, the adapter and the merged checkpoint.

Drives → New drive, Name llm-runs, Kind On each machine's data location, Mounted in workers at /out, Access Read-write, Create drive.

$ curl -fsS -X POST "$API/drives" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' \
    -d '{"metadata": {"name": "llm-runs"}, "spec": {"sources": [{"node_scope": "placed", "path": "", "mount_path": "/out", "mode": "ReadWrite"}]}}' \
    | jq -r .metadata.name
llm-runs

2. Fine-tune with LoRA and merge#

The script trains a LoRA adapter, then merges it into the base model and saves a complete checkpoint (safetensors, with its tokenizer): the format both llama.cpp's converter and vLLM read. The base model and the dataset are cached on the drive, so a second run does not download them again.

finetune.py
import json
import os

import torch
from datasets import load_dataset
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM, AutoTokenizer

BASE = os.environ.get("BASE_MODEL", "Qwen/Qwen2.5-1.5B-Instruct")
OUT = os.environ.get("OUT_DIR", "/out")
STEPS = int(os.environ.get("STEPS", "500"))
BATCH = int(os.environ.get("BATCH_SIZE", "8"))
# ASTRAEUS_JOB_NAME is <namespace>.<run>; keep the run's own name.
RUN = os.environ.get("ASTRAEUS_JOB_NAME", "local").rsplit(".", 1)[-1]
run_dir = os.path.join(OUT, RUN)
os.makedirs(run_dir, exist_ok=True)

tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, torch_dtype=torch.bfloat16).to("cuda")
model.gradient_checkpointing_enable()
model.enable_input_require_grads()
model = get_peft_model(model, LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
))
model.print_trainable_parameters()

rows = load_dataset("yahma/alpaca-cleaned", split="train[:4000]")
def as_chat(r):
    prompt = r["instruction"] + (f"\n\n{r['input']}" if r["input"] else "")
    return tok.apply_chat_template(
        [{"role": "user", "content": prompt}, {"role": "assistant", "content": r["output"]}],
        tokenize=False,
    )
texts = [as_chat(r) for r in rows]

opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=2e-4)
model.train()
for step in range(STEPS):
    i = (step * BATCH) % (len(texts) - BATCH)
    enc = tok(texts[i:i + BATCH], return_tensors="pt", padding=True, truncation=True, max_length=1024).to("cuda")
    labels = enc.input_ids.masked_fill(enc.attention_mask == 0, -100)
    loss = model(**enc, labels=labels).loss
    loss.backward()
    opt.step()
    opt.zero_grad(set_to_none=True)
    if step % 25 == 0:
        print(json.dumps({"step": step, "loss": round(loss.item(), 4)}), flush=True)

model.save_pretrained(os.path.join(run_dir, "adapter"))           # the LoRA adapter alone
merged = model.merge_and_unload()                                  # base + adapter, one model
merged.save_pretrained(os.path.join(run_dir, "merged"), safe_serialization=True)
tok.save_pretrained(os.path.join(run_dir, "merged"))
print(f"done: merged model in {run_dir}/merged", flush=True)
finetune.json
{
  "metadata": {"name": "finetune-qwen", "labels": {"project": "support-bot"}},
  "spec": {
    "start": "Independent",
    "task_template": {
      "image": "pytorch/pytorch:2.4.1-cuda12.4-cudnn9-runtime",
      "command": "sh",
      "args": ["-c", "pip install --no-cache-dir transformers==4.46.3 peft==0.13.2 datasets==3.1.0 accelerate==1.1.1 && python /app/finetune.py"],
      "env": {"OUT_DIR": "/out", "STEPS": "500", "BATCH_SIZE": "8", "HF_HOME": "/out/hf-cache"},
      "restart_policy": "OnFailure",
      "time_limit_seconds": 7200,
      "requested_resources": {
        "cpu_cores": 8,
        "memory_bytes": 34359738368,
        "gpu_requests": {"count": 1, "min_memory_gb": 20},
        "node_selection": {"mode": "Any"}
      },
      "datavolume_refs": [{"name": "llm-runs", "mount_path": "/out", "mode": "ReadWrite"}],
      "configs": []
    }
  }
}
  1. Put the script into the run: jq --rawfile src finetune.py '.spec.task_template.configs = [{"mounts": ["/app/finetune.py"], "value": $src}]' finetune.json > finetune.full.json.
  2. Runs → New run → Edit as JSON, paste finetune.full.json, Start run.
$ jq --rawfile src finetune.py \
    '.spec.task_template.configs = [{"mounts": ["/app/finetune.py"], "value": $src}]' \
    finetune.json > finetune.full.json
$ curl -fsS -X POST "$API/runs" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    -H 'content-type: application/json' -d @finetune.full.json | jq -r .metadata.name
finetune-qwen

Follow it as in the training recipe:

$ astra astraeus logs finetune-qwen-0 -f
trainable params: 18,464,768 || all params: 1,562,179,072 || trainable%: 1.1820
{"step": 0, "loss": …}
{"step": 25, "loss": …}
…
done: merged model in /out/finetune-qwen/merged

The run ends Completed. On the drive llm-runs, finetune-qwen/merged/ holds config.json, model.safetensors (about 3.1 GB in BF16) and the tokenizer; finetune-qwen/adapter/ the adapter alone (about 37 MB).

3. Check that it worked#

  • The run is Completed, and its log ends with done: merged model in /out/finetune-qwen/merged.
  • The loss printed every 25 steps goes down.
  • On llm-runs, finetune-qwen/merged/ holds the merged checkpoint and finetune-qwen/adapter/ the adapter.

Next: serve it#

Serve a fine-tuned model (Eos) takes finetune-qwen/merged from here: converts it to GGUF, registers it as a model from its drive, deploys it and calls it with the OpenAI API. Or serve the merged folder as it is with vLLM (a variation there).

Variations#

Another base model. Any base model with a Hugging Face checkpoint works the same way: change BASE_MODEL, and the GPU memory the run asks for. Bigger models need a bigger GPU, or several (see Distributed training across machines).

Your own data. Put a JSON Lines file with instruction, input and output on the drive (or on a dataset drive) and load it with load_dataset("json", data_files="/out/data.jsonl").

Keep only the adapter. The adapter is about 37 MB; the merged model 3.1 GB. Skip the merge to save space, but merge before serving: Eos serves complete models.

Clean up#

Delete the run, then the drive llm-runs once you have what you need.

Deleting the drive deletes the weights

Drives kept on each machine are the one kind whose files Astraeus deletes: the adapter and the merged checkpoint go with it. Copy out what you want to keep first.

Troubleshooting#

Symptom Cause Fix
torch.OutOfMemoryError in the training log Batch or sequence too large for the GPU. Lower BATCH_SIZE to 4, or max_length to 512.
The run waits Pending No machine has a free GPU with 20 GB. Open Why? next to its state; lower min_memory_gb for a smaller base model.
Downloads fail No outbound HTTPS to Hugging Face or PyPI from the machine. Allow it, or put the base model on the drive first.