Deployment specification#
A deployment serves one model: replicas of an inference engine on your
machines, behind one endpoint, called by the deployment's name. This page
lists every field, its default, the limits the cluster checks, what each
replica is given (the plan) and the states. In the API a deployment is
/deployments in your workspace on a cluster; see
Management API.
{
"metadata": {"name": "chat"},
"spec": {
"model": "llama-3-2-3b",
"engine": "auto",
"replicas": {"min": 1, "max": 1},
"accelerator": "gpu",
"gpus": 1,
"context_length": 8192,
"parallel": 4
}
}
Metadata#
| Field | Type | Default | Description |
|---|---|---|---|
metadata.name |
string | required | Lowercase letters, digits and -, not starting or ending with -, at most 40 characters. It is also what clients put in a request's model. |
metadata.labels |
map | {} |
Your labels. |
spec#
| Field | Type | Default | Description |
|---|---|---|---|
model |
string | required | A model of the same workspace. It must exist; its weights need not be on any machine yet. |
engine |
string | auto |
auto (the model's: GGUF → llama.cpp, safetensors → vllm), llama.cpp or vllm. |
replicas.min |
integer | 1 |
Fewest replicas. 0 scales to zero when idle; the next request starts one. At most max. |
replicas.max |
integer | 1 |
Most replicas, 1 to 64. Between min and max the deployment scales with load. |
replicas.idle_minutes |
integer | 15 |
With min: 0, minutes without a request before the last replica stops, 1 to 1440. Ignored (cleared) when min is above 0. |
accelerator |
string | gpu |
gpu: whole GPUs of one machine per replica. cpu: CPU only, llama.cpp and GGUF only. |
gpus |
integer | chosen | GPUs per replica, all of one machine: 1, 2, 4 or 8. llama.cpp splits the layers across them, vLLM the tensors. When omitted, the fewest that hold the model on the largest-GPU machine your workspace may use, written into the spec when the deployment is created or changed. Not allowed with cpu. |
gpu_vendor |
string | any | nvidia or amd to restrict replicas to that vendor's GPUs. Omitted: any GPU the engine runs on. See GPU vendors. |
cpu_cores |
integer | by size | CPU cores per replica, 1 to 1024. Default with GPUs: 4 per GPU. CPU only: 2 up to 2B parameters, 4 up to 9B, 6 above, at most the cores of the largest machine your workspace may use. |
context_length |
integer | the model's, capped | The longest context of one request, in tokens; at least 256 and never above the model's. Default: llama.cpp the model's, at most 8192; vLLM the model's, at most 32768. |
parallel |
integer | by engine | Requests one replica serves at once, 1 to 1024 (llama.cpp's --parallel, vLLM's --max-num-seqs). Default: llama.cpp on a GPU 4, lowered to 2 or 1 when that is what fits; llama.cpp on a CPU 2; vLLM 256. |
engine_args |
array of strings | [] |
Extra engine flags from a fixed allow-list. See Engine arguments. |
priority |
integer | 0 |
Place in the queue, as a run's. Replicas of a higher-priority deployment start first and may preempt lower-priority work, as runs do (Priorities). |
A CPU-only replica's cores are what it is placed by and its share when the
machine is busy, not a ceiling: it uses the machine's idle cores too, so
several small models share one machine. llama.cpp on a CPU runs with
--threads at twice the replica's cores, at most the machine's.
Engine arguments#
engine_args adds flags after the ones the deployment sets. Rules:
- At most 64 arguments; each one argument, at most 4096 characters, no control characters.
- No shell characters:
`$|&<>\. Arguments are passed to the engine as they are, never to a shell. - Flags in their long form (
--flash-attn, not-fa); the first argument is a flag;--flag valueand--flag=valueboth work. - A flag the deployment sets is refused with where to set it instead.
- Anything not on the allow-list is refused, with the list in the message.
Allowed for llama.cpp: --threads, --threads-batch, --batch-size,
--ubatch-size, --n-predict, --keep, --swa-full, --flash-attn,
--cache-type-k, --cache-type-v, --kv-offload, --no-kv-offload,
--defrag-thold, --rope-scaling, --rope-scale, --rope-freq-base,
--rope-freq-scale, --yarn-orig-ctx, --yarn-ext-factor,
--yarn-attn-factor, --yarn-beta-slow, --yarn-beta-fast,
--override-kv, --cpu-moe, --n-cpu-moe, --seed, --samplers,
--temp, --top-k, --top-p, --min-p, --typical-p,
--repeat-last-n, --repeat-penalty, --presence-penalty,
--frequency-penalty, --ignore-eos, --kv-unified, --no-kv-unified,
--ctx-checkpoints, --cache-ram, --context-shift,
--no-context-shift, --cache-prompt, --no-cache-prompt,
--cache-reuse, --cont-batching, --no-cont-batching, --warmup,
--no-warmup, --pooling, --embeddings, --split-mode,
--tensor-split, --main-gpu, --reranking, --jinja, --no-jinja,
--reasoning-format, --reasoning-effort, --reasoning-budget,
--chat-template-kwargs, --timeout, --threads-http,
--sse-ping-interval, --slots, --no-slots, --fit, --fit-target,
--fit-ctx.
Allowed for vLLM: --dtype, --quantization, --kv-cache-dtype,
--enable-prefix-caching, --no-enable-prefix-caching,
--enforce-eager, --no-enforce-eager, --max-num-batched-tokens,
--enable-chunked-prefill, --no-enable-chunked-prefill, --block-size,
--cpu-offload-gb, --kv-cache-memory-bytes, --gpu-memory-utilization,
--seed, --generation-config, --override-generation-config,
--max-logprobs, --logprobs-mode, --disable-log-stats,
--limit-mm-per-prompt, --tokenizer-mode, --hf-overrides,
--scheduling-policy, --long-prefill-token-threshold,
--compilation-config, --async-scheduling, --no-async-scheduling,
--enable-auto-tool-choice, --tool-call-parser, --reasoning-parser,
--chat-template-content-format, --default-chat-template-kwargs,
--response-role, --enable-force-include-usage,
--enable-prompt-tokens-details, --return-tokens-as-token-ids,
--max-log-len, --disable-uvicorn-access-log, --uvicorn-log-level,
--enable-server-load-tracking.
Not allowed, by design: flags that read a file or fetch from the network
(adapters, grammars, draft models), keys and certificates, and switches
that open the server further (--trust-remote-code, --api-key).
Set by the deployment (refused in engine_args):
| Engine | Flag | Set it with |
|---|---|---|
| both | --model, --host, --port |
the deployment |
| llama.cpp | --ctx-size |
spec.context_length |
| llama.cpp | --parallel |
spec.parallel |
| llama.cpp | --n-gpu-layers, --gpu-layers |
spec.accelerator |
| llama.cpp | --alias |
the deployment's name |
| llama.cpp | --metrics, --webui, --no-webui |
the deployment |
| llama.cpp | --chat-template, --chat-template-file |
the model's chat_template |
| llama.cpp | --mmproj |
the model's mmproj |
| vLLM | --served-model-name |
the deployment's name |
| vLLM | --max-model-len |
spec.context_length |
| vLLM | --max-num-seqs |
spec.parallel |
| vLLM | --tensor-parallel-size, --pipeline-parallel-size |
spec.gpus (a replica is one machine's GPUs) |
| vLLM | --runner, --convert |
the model's capabilities |
| vLLM | --chat-template |
the model's chat_template |
What each replica runs#
llama.cpp (release v0.5.0) |
vLLM (release v0.30.0) |
|
|---|---|---|
| Image on NVIDIA | ghcr.io/ggml-org/llama.cpp:server-cuda-v0.5.0 |
vllm/vllm-openai:v0.30.0 |
| Image on AMD | the ROCm build; the Vulkan build for Radeon GPUs ROCm has no kernels for | vllm/vllm-openai-rocm:v0.30.0 |
| CPU only | the CPU build | — |
| Mac | the native Metal build (b11146) |
— |
| Weights | /model/<file>.gguf |
/model |
Name in model |
--alias <deployment> |
--served-model-name <deployment> |
| Context | --ctx-size <context × parallel> |
--max-model-len <context> |
| At once | --parallel <n> |
--max-num-seqs <n> |
| On GPUs | --n-gpu-layers all; --split-mode layer with more than one |
--gpu-memory-utilization 0.90; --tensor-parallel-size <n> with more than one |
| Embeddings | --embeddings [--pooling <p>] |
--runner pooling --convert embed |
| Tools | --jinja (the model's template) |
--enable-auto-tool-choice --tool-call-parser <p> |
| Reasoning | the template's | --reasoning-parser <p> |
| Always | --metrics --jinja --no-webui |
HF_HUB_OFFLINE=1, VLLM_NO_USAGE_STATS=1, DO_NOT_TRACK=1 |
Images are pinned by digest: what serves a model does not change under it.
llama.cpp's --ctx-size is the whole cache, split between its parallel
slots, so each request gets the whole context_length.
Every replica also mounts:
- the model's weights, read-only, at
/model; - the workspace's engine cache, a drive named
engine-cachekept on each machine, at/root/.cache: compiled kernels and caches that make the second start of a replica quicker.
Each replica listens on port 8000. It takes traffic once GET /health
answers (checked every 5 s, 3 s timeout, 3 failures in a row to be taken
out). Engine metrics are read from GET /metrics every 15 s.
Memory per replica#
| Engine | Host memory per replica |
|---|---|
| llama.cpp on GPUs | weights + 2 GiB |
| llama.cpp on a CPU | weights + 2 GiB, or the memory estimate + 1 GiB, whichever is larger |
| vLLM | weights + 8 GiB, + 2 GiB per GPU beyond the first |
At least 2 GiB. Each GPU of a replica must have at least the plan's
min_gpu_memory_gb; see Fit and quantization.
Scaling#
replicas |
Behaviour |
|---|---|
min = max |
That many replicas, always. |
min < max |
Scales on requests in flight per replica, against three quarters of parallel (llama.cpp's llamacpp:requests_processing, vLLM's vllm:num_requests_running), at most once every 120 s. |
min: 0 |
After idle_minutes without a request the last replica stops (ScaledToZero). The next request starts one replica and waits for it. |
What the cluster returns#
GET /deployments/{name} returns the specification (with gpus and, for
llama.cpp, parallel filled in when they were chosen) and:
| Field | Description |
|---|---|
status.state, status.reason |
See States. |
replicas, ready_replicas |
Replicas wanted now, and serving. |
members[] |
Each replica: run, worker, machine, state, reason, ready (behind the endpoint, health check passing). |
endpoint.name |
deploy-<name>. |
endpoint.internal |
deploy-<name>.<namespace>.astraeus.local:8000: the address your runs on the cluster's machines can call. |
endpoint.listen_port, endpoint.external |
The port on machines that serve external access, and <host>:<port> when one does. |
kind |
chat or embedding. |
served_model_name |
What a client puts in model: the deployment's name. |
plan |
What each replica runs and takes: engine, image, amd_images, gpu_vendor, accelerator, kind, gpus, min_gpu_memory_gb, memory_basis (estimate, arch, catalog or weights), cpu_cores, cpu_burst, threads, memory_bytes, context_length, parallel, weights_bytes, port. |
gateways[] |
The base URLs (without /v1) where programs call it with an API key. Empty when no machine of yours runs the gateway. |
capabilities |
What it answers: chat, embeddings, vision, audio, video, transcription, tools (and tools_reason, tools_engine_args when not), reasoning, reasoning_switch (enable_thinking or reasoning_effort), structured_output, logprobs, engine, tool_call_parser, reasoning_parser. |
States#
| State | Meaning |
|---|---|
Pending |
Waiting for a machine, or for the weights to be downloaded. The reason is the furthest replica's. |
Starting |
A replica is starting: its weights are being fetched or the engine is loading them. |
Ready |
At least one replica serves (1 replica serving, 1 of 3 replicas serving). |
ScaledToZero |
No replica runs; the next request starts one. |
Failed |
It cannot be served as asked: the model is gone or its download failed, the specification does not fit the model, or a replica keeps exiting (3 restarts without serving). The reason says which. |
Changes and deletion#
PUT /deployments/{name}with a newspecchanges it. The cluster checks the new specification against the model first. Replicas are then replaced one at a time.- Deleting a deployment stops its replicas and removes its endpoint and its shares. The model and its weights stay.
Rules#
| Rule | Message |
|---|---|
| A model of the same workspace | spec.model "x": a model of the deployment's own workspace |
replicas.max between 1 and 64 |
spec.replicas.max: between 1 and 64 |
replicas.min ≤ max |
spec.replicas.min: at most spec.replicas.max |
idle_minutes 1 to 1440 |
spec.replicas.idle_minutes: between 1 and 1440 |
context_length ≥ 256 |
spec.context_length: at least 256 tokens |
context_length ≤ the model's |
spec.context_length 200000: model llama-3-2-3b takes at most 131072 tokens |
parallel 1 to 1024 |
spec.parallel: between 1 and 1024 |
cpu_cores 1 to 1024 |
spec.cpu_cores: between 1 and 1024 |
gpus 1, 2, 4, 8 |
spec.gpus: 1, 2, 4 or 8 GPUs of one machine per replica |
| No GPUs on a CPU deployment | spec.gpus: a CPU-only deployment has none |
| vLLM needs a GPU | spec.accelerator: cpu is served by llama.cpp; vLLM needs a GPU |
| The engine matches the format | llama.cpp serves GGUF; model x is a safetensors checkpoint: serve it with vllm, model x is GGUF: serve it with llama.cpp (vLLM serves safetensors checkpoints here) |
Invalid specifications are refused with 400 INVALID_DEPLOYMENT; see
Errors and limits.