Skip to content

Model specification#

A model is a set of weights and how to serve them. This page lists every field of a model, what the cluster fills in, the rules it checks and the states a model goes through. In the API a model is /models in your workspace on a cluster; see Management API.

model.json
{
  "metadata": {"name": "llama-3-2-3b"},
  "spec": {
    "source": {
      "huggingface": {
        "repo": "bartowski/Llama-3.2-3B-Instruct-GGUF",
        "revision": "5ab33fa94d1d04e903623ae72c95d1696f09f9e8",
        "files": [
          {
            "path": "Llama-3.2-3B-Instruct-Q4_K_M.gguf",
            "size_bytes": 2019377696,
            "sha256": "6c1a2b41161032677be168d354123594c0e6e67d2b9227c84f296ad037c728ff"
          }
        ]
      }
    },
    "display": "Llama 3.2 3B Instruct",
    "params_b": 3.2,
    "quantization": "Q4_K_M",
    "context_length": 131072
  }
}

Metadata#

Field Type Default Description
metadata.name string required Lowercase letters, digits and -, not starting or ending with -, at most 40 characters. Local to your workspace.
metadata.labels map {} Your labels.

spec#

Field Type Default Description
source object required Where the weights come from: exactly one of huggingface or drive. See Source.
format string from the files gguf (one file, or split parts, for llama.cpp) or safetensors (a Hugging Face checkpoint, for vLLM). Derived from the file names of a Hugging Face source; required for a drive source.
engine string from the format llama.cpp or vllm. GGUF defaults to llama.cpp, safetensors to vllm. llama.cpp cannot serve safetensors.
display string "" What people call it (Llama 3.2 3B Instruct).
params_b number none Parameters, in billions. Positive. Used for CPU sizing and fit when known.
quantization string "" A label: Q4_K_M, Q8_0, FP8, BF16… Informational.
context_length integer none The longest context the model takes, in tokens. A deployment cannot ask for more. When a machine has read a GGUF file's header, the header's value is used instead.
chat_template string "" A Jinja chat template to use instead of the model's own. At most 64 KiB. Written into each replica and passed to the engine.
catalog string "" <name>:<tag> when the model came from the model library. Set by the library.
min_gpu_memory_gb number none GPU memory needed to serve it, GB, as the library says: used until a machine has read the weights' header. Between 0 and 4096.
capabilities array [] (chat) What the model does: chat, vision, embedding, tools, reasoning, audio, video, transcription. See Capabilities.
hints object {} How the engines serve it. See Engine hints.
mmproj string "" A vision or audio GGUF's projector file (llama.cpp's --mmproj). Must end in .gguf and be one of the model's files. GGUF only.
arch object none The model's shape, for the memory estimate before any machine has read the weights. See Shape.

Source#

From Hugging Face — source.huggingface. The weights are fetched from the hub at a pinned commit into a drive the model manages.

Field Type Default Description
repo string required owner/name. Each part: letters, digits, ., _, -, at most 96 characters, not starting with . or -.
revision string required A commit: 40 hexadecimal characters. A tag or branch is refused, because it can move.
files array required At least one. The files to fetch.
files[].path string required The path in the repository. Relative, at most 512 characters, letters, digits and ._-/+, no .. and no part starting with .. Listed once.
files[].size_bytes integer required The file's size. Not 0.
files[].sha256 string "" The file's SHA-256 (64 hexadecimal characters). When given, the download is checked against it and a mismatch is removed and fails.
credential string "" A credential with a token key, for a gated or private repository. The machine that downloads resolves it; Astralyx stores only its name.

From a drive — source.drive. The weights are already in one of your drives, and are used where they are.

Field Type Default Description
name string required The drive.
path string "" (the drive's root) A folder in the drive, relative, without ... For llama.cpp, the path names the .gguf file itself (the first part of a split model), for example merged/model-q4_k_m.gguf. For vLLM, the folder that holds the checkpoint (config.json, *.safetensors, the tokenizer).

Capabilities#

Value What it means
chat Conversation: /v1/chat/completions, /v1/completions. The default when the list is empty.
vision Images in a conversation (image_url parts). With llama.cpp only when mmproj is set.
embedding Vectors: the deployment serves /v1/embeddings and nothing else.
tools Calls tools. llama.cpp always runs with the model's template (--jinja); vLLM needs a tool-call parser (hints.tool_call_parser).
reasoning Thinks before it answers. With vLLM, hints.reasoning_parser separates the thinking.
audio Audio in a conversation (input_audio parts). vLLM, or llama.cpp with mmproj.
video Short videos in a conversation (video_url parts). vLLM only.
transcription Speech to text: /v1/audio/transcriptions. vLLM only. A model with transcription and without chat (Whisper) is served for transcription only.

Engine hints#

Field Type Default Description
hints.tool_call_parser string "" vLLM's --tool-call-parser (hermes, llama3_json…). With tools, the deployment adds --enable-auto-tool-choice --tool-call-parser <it>.
hints.reasoning_parser string "" vLLM's --reasoning-parser (deepseek_r1, qwen3…), added with reasoning.
hints.pooling string "" llama.cpp's --pooling for an embedding model: none, mean, cls, last or rank.

Parser names: letters, digits, _, -, ., at most 64 characters.

Shape#

Field Type Description
arch.architecture string llama, qwen2… At most 64 characters.
arch.layers integer Layers (at most 1024).
arch.kv_heads integer Key/value heads (at most 1024).
arch.head_dim integer Each head's dimension (at most 4096).
arch.moe_active_params_b number A mixture of experts: parameters active per token, billions.

What the cluster returns#

GET /models/{name} returns the specification with format and engine filled in, and:

Field Description
status.state, status.reason See States.
drive The drive the weights are in: model-<name> for a Hugging Face model, or your drive.
size_bytes The listed files' total size (0 for a drive source).
copies[] machine, bytes, complete: each machine's copy of the weights.
pulls[] machine, run, state, reason: each download run.
gguf The GGUF header, once a machine with a whole copy has read it.
memory_estimate_bytes, estimate_context Memory to serve it (weights, KV cache, overhead) at estimate_context tokens.

States#

State Reason (examples) Meaning
Registered Not on any machine yet No copy yet. The first deployment's replica, or a pull, downloads it.
Pulling Pulling onto gpu-01 A download run is in progress.
Ready On gpu-01, On 2 machines, Weights in drive llm-runs At least one whole copy; for a drive source, while the drive exists.
Failed Pulling onto gpu-01 failed: …, Drive llm-runs does not exist A download failed and no copy is whole, or the drive is gone.

The weights of a Hugging Face model#

Creating a model from Hugging Face also creates its drive, model-<name>: a drive kept on each machine's data location, sized by the files, that holds a cache of the hub. A download is a run named fill-model-<name>-<machine>. It fetches each file at the pinned revision with resume, checks its SHA-256 when given, and marks the copy whole last. Its log prints progress every ten seconds. The machine's copy is at <data location>/drives/<namespace>.model-<name>/.

  • The drive cannot be deleted on its own; it goes with the model.
  • Its copies can be evicted when a machine needs room: the hub has the files.
  • A drive named model-<name> that is yours blocks a model of that name (409 MODEL_CONFLICT).

Rules#

Rule Error
Exactly one of huggingface or drive 400 INVALID_MODEL spec.source: exactly one of huggingface or drive
revision is a commit spec.source.huggingface.revision: a commit (40 hex characters) — a tag or branch can move
A drive source says its format spec.format: gguf or safetensors (weights in a drive do not say)
The files are all GGUF, or a safetensors checkpoint spec.format: the files are neither all GGUF nor a safetensors checkpoint; say which
llama.cpp serves GGUF only spec.engine: llama.cpp serves GGUF; a safetensors checkpoint is served by vllm
The name is free 409 MODEL_ALREADY_EXISTS
Not deleted while used 409 MODEL_IN_USE model llama-3-2-3b is used by 2 live worker(s); add ?force=true

Deleting a Hugging Face model deletes its drive — every machine's copy — and stops its downloads. Deleting a model whose weights are in your drive never touches the drive. Deleting a model does not delete its deployments; they fail with the model gone.