Model specification
A model is a set of weights and how to serve them. This page lists
every field of a model, what the cluster fills in, the rules it checks and
the states a model goes through. In the API a model is
/models in your workspace on a cluster; see Management API.
model.json{
"metadata": {"name": "llama-3-2-3b"},
"spec": {
"source": {
"huggingface": {
"repo": "bartowski/Llama-3.2-3B-Instruct-GGUF",
"revision": "5ab33fa94d1d04e903623ae72c95d1696f09f9e8",
"files": [
{
"path": "Llama-3.2-3B-Instruct-Q4_K_M.gguf",
"size_bytes": 2019377696,
"sha256": "6c1a2b41161032677be168d354123594c0e6e67d2b9227c84f296ad037c728ff"
}
]
}
},
"display": "Llama 3.2 3B Instruct",
"params_b": 3.2,
"quantization": "Q4_K_M",
"context_length": 131072
}
}
| Field |
Type |
Default |
Description |
metadata.name |
string |
required |
Lowercase letters, digits and -, not starting or ending with -, at most 40 characters. Local to your workspace. |
metadata.labels |
map |
{} |
Your labels. |
spec
| Field |
Type |
Default |
Description |
source |
object |
required |
Where the weights come from: exactly one of huggingface or drive. See Source. |
format |
string |
from the files |
gguf (one file, or split parts, for llama.cpp) or safetensors (a Hugging Face checkpoint, for vLLM). Derived from the file names of a Hugging Face source; required for a drive source. |
engine |
string |
from the format |
llama.cpp or vllm. GGUF defaults to llama.cpp, safetensors to vllm. llama.cpp cannot serve safetensors. |
display |
string |
"" |
What people call it (Llama 3.2 3B Instruct). |
params_b |
number |
none |
Parameters, in billions. Positive. Used for CPU sizing and fit when known. |
quantization |
string |
"" |
A label: Q4_K_M, Q8_0, FP8, BF16… Informational. |
context_length |
integer |
none |
The longest context the model takes, in tokens. A deployment cannot ask for more. When a machine has read a GGUF file's header, the header's value is used instead. |
chat_template |
string |
"" |
A Jinja chat template to use instead of the model's own. At most 64 KiB. Written into each replica and passed to the engine. |
catalog |
string |
"" |
<name>:<tag> when the model came from the model library. Set by the library. |
min_gpu_memory_gb |
number |
none |
GPU memory needed to serve it, GB, as the library says: used until a machine has read the weights' header. Between 0 and 4096. |
capabilities |
array |
[] (chat) |
What the model does: chat, vision, embedding, tools, reasoning, audio, video, transcription. See Capabilities. |
hints |
object |
{} |
How the engines serve it. See Engine hints. |
mmproj |
string |
"" |
A vision or audio GGUF's projector file (llama.cpp's --mmproj). Must end in .gguf and be one of the model's files. GGUF only. |
arch |
object |
none |
The model's shape, for the memory estimate before any machine has read the weights. See Shape. |
Source
From Hugging Face — source.huggingface. The weights are fetched from
the hub at a pinned commit into a drive the model manages.
| Field |
Type |
Default |
Description |
repo |
string |
required |
owner/name. Each part: letters, digits, ., _, -, at most 96 characters, not starting with . or -. |
revision |
string |
required |
A commit: 40 hexadecimal characters. A tag or branch is refused, because it can move. |
files |
array |
required |
At least one. The files to fetch. |
files[].path |
string |
required |
The path in the repository. Relative, at most 512 characters, letters, digits and ._-/+, no .. and no part starting with .. Listed once. |
files[].size_bytes |
integer |
required |
The file's size. Not 0. |
files[].sha256 |
string |
"" |
The file's SHA-256 (64 hexadecimal characters). When given, the download is checked against it and a mismatch is removed and fails. |
credential |
string |
"" |
A credential with a token key, for a gated or private repository. The machine that downloads resolves it; Astralyx stores only its name. |
From a drive — source.drive. The weights are already in one of your
drives, and are used where they are.
| Field |
Type |
Default |
Description |
name |
string |
required |
The drive. |
path |
string |
"" (the drive's root) |
A folder in the drive, relative, without ... For llama.cpp, the path names the .gguf file itself (the first part of a split model), for example merged/model-q4_k_m.gguf. For vLLM, the folder that holds the checkpoint (config.json, *.safetensors, the tokenizer). |
Capabilities
| Value |
What it means |
chat |
Conversation: /v1/chat/completions, /v1/completions. The default when the list is empty. |
vision |
Images in a conversation (image_url parts). With llama.cpp only when mmproj is set. |
embedding |
Vectors: the deployment serves /v1/embeddings and nothing else. |
tools |
Calls tools. llama.cpp always runs with the model's template (--jinja); vLLM needs a tool-call parser (hints.tool_call_parser). |
reasoning |
Thinks before it answers. With vLLM, hints.reasoning_parser separates the thinking. |
audio |
Audio in a conversation (input_audio parts). vLLM, or llama.cpp with mmproj. |
video |
Short videos in a conversation (video_url parts). vLLM only. |
transcription |
Speech to text: /v1/audio/transcriptions. vLLM only. A model with transcription and without chat (Whisper) is served for transcription only. |
Engine hints
| Field |
Type |
Default |
Description |
hints.tool_call_parser |
string |
"" |
vLLM's --tool-call-parser (hermes, llama3_json…). With tools, the deployment adds --enable-auto-tool-choice --tool-call-parser <it>. |
hints.reasoning_parser |
string |
"" |
vLLM's --reasoning-parser (deepseek_r1, qwen3…), added with reasoning. |
hints.pooling |
string |
"" |
llama.cpp's --pooling for an embedding model: none, mean, cls, last or rank. |
Parser names: letters, digits, _, -, ., at most 64 characters.
Shape
| Field |
Type |
Description |
arch.architecture |
string |
llama, qwen2… At most 64 characters. |
arch.layers |
integer |
Layers (at most 1024). |
arch.kv_heads |
integer |
Key/value heads (at most 1024). |
arch.head_dim |
integer |
Each head's dimension (at most 4096). |
arch.moe_active_params_b |
number |
A mixture of experts: parameters active per token, billions. |
What the cluster returns
GET /models/{name} returns the specification with format and engine
filled in, and:
| Field |
Description |
status.state, status.reason |
See States. |
drive |
The drive the weights are in: model-<name> for a Hugging Face model, or your drive. |
size_bytes |
The listed files' total size (0 for a drive source). |
copies[] |
machine, bytes, complete: each machine's copy of the weights. |
pulls[] |
machine, run, state, reason: each download run. |
gguf |
The GGUF header, once a machine with a whole copy has read it. |
memory_estimate_bytes, estimate_context |
Memory to serve it (weights, KV cache, overhead) at estimate_context tokens. |
States
| State |
Reason (examples) |
Meaning |
Registered |
Not on any machine yet |
No copy yet. The first deployment's replica, or a pull, downloads it. |
Pulling |
Pulling onto gpu-01 |
A download run is in progress. |
Ready |
On gpu-01, On 2 machines, Weights in drive llm-runs |
At least one whole copy; for a drive source, while the drive exists. |
Failed |
Pulling onto gpu-01 failed: …, Drive llm-runs does not exist |
A download failed and no copy is whole, or the drive is gone. |
The weights of a Hugging Face model
Creating a model from Hugging Face also creates its drive, model-<name>:
a drive kept on each machine's data location, sized by the files, that
holds a cache of the hub. A download is a run named
fill-model-<name>-<machine>. It fetches each file at the pinned revision
with resume, checks its SHA-256 when given, and marks the copy whole last.
Its log prints progress every ten seconds. The machine's copy is at
<data location>/drives/<namespace>.model-<name>/.
- The drive cannot be deleted on its own; it goes with the model.
- Its copies can be evicted when a machine needs room: the hub has the files.
- A drive named
model-<name> that is yours blocks a model of that name
(409 MODEL_CONFLICT).
Rules
| Rule |
Error |
Exactly one of huggingface or drive |
400 INVALID_MODEL spec.source: exactly one of huggingface or drive |
revision is a commit |
spec.source.huggingface.revision: a commit (40 hex characters) — a tag or branch can move |
| A drive source says its format |
spec.format: gguf or safetensors (weights in a drive do not say) |
| The files are all GGUF, or a safetensors checkpoint |
spec.format: the files are neither all GGUF nor a safetensors checkpoint; say which |
| llama.cpp serves GGUF only |
spec.engine: llama.cpp serves GGUF; a safetensors checkpoint is served by vllm |
| The name is free |
409 MODEL_ALREADY_EXISTS |
| Not deleted while used |
409 MODEL_IN_USE model llama-3-2-3b is used by 2 live worker(s); add ?force=true |
Deleting a Hugging Face model deletes its drive — every machine's copy —
and stops its downloads. Deleting a model whose weights are in your drive
never touches the drive. Deleting a model does not delete its deployments;
they fail with the model gone.