Skip to content

Deployment specification#

A deployment serves one model: replicas of an inference engine on your machines, behind one endpoint, called by the deployment's name. This page lists every field, its default, the limits the cluster checks, what each replica is given (the plan) and the states. In the API a deployment is /deployments in your workspace on a cluster; see Management API.

deployment.json
{
  "metadata": {"name": "chat"},
  "spec": {
    "model": "llama-3-2-3b",
    "engine": "auto",
    "replicas": {"min": 1, "max": 1},
    "accelerator": "gpu",
    "gpus": 1,
    "context_length": 8192,
    "parallel": 4
  }
}

Metadata#

Field Type Default Description
metadata.name string required Lowercase letters, digits and -, not starting or ending with -, at most 40 characters. It is also what clients put in a request's model.
metadata.labels map {} Your labels.

spec#

Field Type Default Description
model string required A model of the same workspace. It must exist; its weights need not be on any machine yet.
engine string auto auto (the model's: GGUF → llama.cpp, safetensors → vllm), llama.cpp or vllm.
replicas.min integer 1 Fewest replicas. 0 scales to zero when idle; the next request starts one. At most max.
replicas.max integer 1 Most replicas, 1 to 64. Between min and max the deployment scales with load.
replicas.idle_minutes integer 15 With min: 0, minutes without a request before the last replica stops, 1 to 1440. Ignored (cleared) when min is above 0.
accelerator string gpu gpu: whole GPUs of one machine per replica. cpu: CPU only, llama.cpp and GGUF only.
gpus integer chosen GPUs per replica, all of one machine: 1, 2, 4 or 8. llama.cpp splits the layers across them, vLLM the tensors. When omitted, the fewest that hold the model on the largest-GPU machine your workspace may use, written into the spec when the deployment is created or changed. Not allowed with cpu.
gpu_vendor string any nvidia or amd to restrict replicas to that vendor's GPUs. Omitted: any GPU the engine runs on. See GPU vendors.
cpu_cores integer by size CPU cores per replica, 1 to 1024. Default with GPUs: 4 per GPU. CPU only: 2 up to 2B parameters, 4 up to 9B, 6 above, at most the cores of the largest machine your workspace may use.
context_length integer the model's, capped The longest context of one request, in tokens; at least 256 and never above the model's. Default: llama.cpp the model's, at most 8192; vLLM the model's, at most 32768.
parallel integer by engine Requests one replica serves at once, 1 to 1024 (llama.cpp's --parallel, vLLM's --max-num-seqs). Default: llama.cpp on a GPU 4, lowered to 2 or 1 when that is what fits; llama.cpp on a CPU 2; vLLM 256.
engine_args array of strings [] Extra engine flags from a fixed allow-list. See Engine arguments.
priority integer 0 Place in the queue, as a run's. Replicas of a higher-priority deployment start first and may preempt lower-priority work, as runs do (Priorities).

A CPU-only replica's cores are what it is placed by and its share when the machine is busy, not a ceiling: it uses the machine's idle cores too, so several small models share one machine. llama.cpp on a CPU runs with --threads at twice the replica's cores, at most the machine's.

Engine arguments#

engine_args adds flags after the ones the deployment sets. Rules:

  • At most 64 arguments; each one argument, at most 4096 characters, no control characters.
  • No shell characters: ` $ | & < > \. Arguments are passed to the engine as they are, never to a shell.
  • Flags in their long form (--flash-attn, not -fa); the first argument is a flag; --flag value and --flag=value both work.
  • A flag the deployment sets is refused with where to set it instead.
  • Anything not on the allow-list is refused, with the list in the message.

Allowed for llama.cpp: --threads, --threads-batch, --batch-size, --ubatch-size, --n-predict, --keep, --swa-full, --flash-attn, --cache-type-k, --cache-type-v, --kv-offload, --no-kv-offload, --defrag-thold, --rope-scaling, --rope-scale, --rope-freq-base, --rope-freq-scale, --yarn-orig-ctx, --yarn-ext-factor, --yarn-attn-factor, --yarn-beta-slow, --yarn-beta-fast, --override-kv, --cpu-moe, --n-cpu-moe, --seed, --samplers, --temp, --top-k, --top-p, --min-p, --typical-p, --repeat-last-n, --repeat-penalty, --presence-penalty, --frequency-penalty, --ignore-eos, --kv-unified, --no-kv-unified, --ctx-checkpoints, --cache-ram, --context-shift, --no-context-shift, --cache-prompt, --no-cache-prompt, --cache-reuse, --cont-batching, --no-cont-batching, --warmup, --no-warmup, --pooling, --embeddings, --split-mode, --tensor-split, --main-gpu, --reranking, --jinja, --no-jinja, --reasoning-format, --reasoning-effort, --reasoning-budget, --chat-template-kwargs, --timeout, --threads-http, --sse-ping-interval, --slots, --no-slots, --fit, --fit-target, --fit-ctx.

Allowed for vLLM: --dtype, --quantization, --kv-cache-dtype, --enable-prefix-caching, --no-enable-prefix-caching, --enforce-eager, --no-enforce-eager, --max-num-batched-tokens, --enable-chunked-prefill, --no-enable-chunked-prefill, --block-size, --cpu-offload-gb, --kv-cache-memory-bytes, --gpu-memory-utilization, --seed, --generation-config, --override-generation-config, --max-logprobs, --logprobs-mode, --disable-log-stats, --limit-mm-per-prompt, --tokenizer-mode, --hf-overrides, --scheduling-policy, --long-prefill-token-threshold, --compilation-config, --async-scheduling, --no-async-scheduling, --enable-auto-tool-choice, --tool-call-parser, --reasoning-parser, --chat-template-content-format, --default-chat-template-kwargs, --response-role, --enable-force-include-usage, --enable-prompt-tokens-details, --return-tokens-as-token-ids, --max-log-len, --disable-uvicorn-access-log, --uvicorn-log-level, --enable-server-load-tracking.

Not allowed, by design: flags that read a file or fetch from the network (adapters, grammars, draft models), keys and certificates, and switches that open the server further (--trust-remote-code, --api-key).

Set by the deployment (refused in engine_args):

Engine Flag Set it with
both --model, --host, --port the deployment
llama.cpp --ctx-size spec.context_length
llama.cpp --parallel spec.parallel
llama.cpp --n-gpu-layers, --gpu-layers spec.accelerator
llama.cpp --alias the deployment's name
llama.cpp --metrics, --webui, --no-webui the deployment
llama.cpp --chat-template, --chat-template-file the model's chat_template
llama.cpp --mmproj the model's mmproj
vLLM --served-model-name the deployment's name
vLLM --max-model-len spec.context_length
vLLM --max-num-seqs spec.parallel
vLLM --tensor-parallel-size, --pipeline-parallel-size spec.gpus (a replica is one machine's GPUs)
vLLM --runner, --convert the model's capabilities
vLLM --chat-template the model's chat_template

What each replica runs#

llama.cpp (release v0.5.0) vLLM (release v0.30.0)
Image on NVIDIA ghcr.io/ggml-org/llama.cpp:server-cuda-v0.5.0 vllm/vllm-openai:v0.30.0
Image on AMD the ROCm build; the Vulkan build for Radeon GPUs ROCm has no kernels for vllm/vllm-openai-rocm:v0.30.0
CPU only the CPU build —
Mac the native Metal build (b11146) —
Weights /model/<file>.gguf /model
Name in model --alias <deployment> --served-model-name <deployment>
Context --ctx-size <context × parallel> --max-model-len <context>
At once --parallel <n> --max-num-seqs <n>
On GPUs --n-gpu-layers all; --split-mode layer with more than one --gpu-memory-utilization 0.90; --tensor-parallel-size <n> with more than one
Embeddings --embeddings [--pooling <p>] --runner pooling --convert embed
Tools --jinja (the model's template) --enable-auto-tool-choice --tool-call-parser <p>
Reasoning the template's --reasoning-parser <p>
Always --metrics --jinja --no-webui HF_HUB_OFFLINE=1, VLLM_NO_USAGE_STATS=1, DO_NOT_TRACK=1

Images are pinned by digest: what serves a model does not change under it. llama.cpp's --ctx-size is the whole cache, split between its parallel slots, so each request gets the whole context_length.

Every replica also mounts:

  • the model's weights, read-only, at /model;
  • the workspace's engine cache, a drive named engine-cache kept on each machine, at /root/.cache: compiled kernels and caches that make the second start of a replica quicker.

Each replica listens on port 8000. It takes traffic once GET /health answers (checked every 5 s, 3 s timeout, 3 failures in a row to be taken out). Engine metrics are read from GET /metrics every 15 s.

Memory per replica#

Engine Host memory per replica
llama.cpp on GPUs weights + 2 GiB
llama.cpp on a CPU weights + 2 GiB, or the memory estimate + 1 GiB, whichever is larger
vLLM weights + 8 GiB, + 2 GiB per GPU beyond the first

At least 2 GiB. Each GPU of a replica must have at least the plan's min_gpu_memory_gb; see Fit and quantization.

Scaling#

replicas Behaviour
min = max That many replicas, always.
min < max Scales on requests in flight per replica, against three quarters of parallel (llama.cpp's llamacpp:requests_processing, vLLM's vllm:num_requests_running), at most once every 120 s.
min: 0 After idle_minutes without a request the last replica stops (ScaledToZero). The next request starts one replica and waits for it.

What the cluster returns#

GET /deployments/{name} returns the specification (with gpus and, for llama.cpp, parallel filled in when they were chosen) and:

Field Description
status.state, status.reason See States.
replicas, ready_replicas Replicas wanted now, and serving.
members[] Each replica: run, worker, machine, state, reason, ready (behind the endpoint, health check passing).
endpoint.name deploy-<name>.
endpoint.internal deploy-<name>.<namespace>.astraeus.local:8000: the address your runs on the cluster's machines can call.
endpoint.listen_port, endpoint.external The port on machines that serve external access, and <host>:<port> when one does.
kind chat or embedding.
served_model_name What a client puts in model: the deployment's name.
plan What each replica runs and takes: engine, image, amd_images, gpu_vendor, accelerator, kind, gpus, min_gpu_memory_gb, memory_basis (estimate, arch, catalog or weights), cpu_cores, cpu_burst, threads, memory_bytes, context_length, parallel, weights_bytes, port.
gateways[] The base URLs (without /v1) where programs call it with an API key. Empty when no machine of yours runs the gateway.
capabilities What it answers: chat, embeddings, vision, audio, video, transcription, tools (and tools_reason, tools_engine_args when not), reasoning, reasoning_switch (enable_thinking or reasoning_effort), structured_output, logprobs, engine, tool_call_parser, reasoning_parser.

States#

State Meaning
Pending Waiting for a machine, or for the weights to be downloaded. The reason is the furthest replica's.
Starting A replica is starting: its weights are being fetched or the engine is loading them.
Ready At least one replica serves (1 replica serving, 1 of 3 replicas serving).
ScaledToZero No replica runs; the next request starts one.
Failed It cannot be served as asked: the model is gone or its download failed, the specification does not fit the model, or a replica keeps exiting (3 restarts without serving). The reason says which.

Changes and deletion#

  • PUT /deployments/{name} with a new spec changes it. The cluster checks the new specification against the model first. Replicas are then replaced one at a time.
  • Deleting a deployment stops its replicas and removes its endpoint and its shares. The model and its weights stay.

Rules#

Rule Message
A model of the same workspace spec.model "x": a model of the deployment's own workspace
replicas.max between 1 and 64 spec.replicas.max: between 1 and 64
replicas.min ≤ max spec.replicas.min: at most spec.replicas.max
idle_minutes 1 to 1440 spec.replicas.idle_minutes: between 1 and 1440
context_length ≥ 256 spec.context_length: at least 256 tokens
context_length ≤ the model's spec.context_length 200000: model llama-3-2-3b takes at most 131072 tokens
parallel 1 to 1024 spec.parallel: between 1 and 1024
cpu_cores 1 to 1024 spec.cpu_cores: between 1 and 1024
gpus 1, 2, 4, 8 spec.gpus: 1, 2, 4 or 8 GPUs of one machine per replica
No GPUs on a CPU deployment spec.gpus: a CPU-only deployment has none
vLLM needs a GPU spec.accelerator: cpu is served by llama.cpp; vLLM needs a GPU
The engine matches the format llama.cpp serves GGUF; model x is a safetensors checkpoint: serve it with vllm, model x is GGUF: serve it with llama.cpp (vLLM serves safetensors checkpoints here)

Invalid specifications are refused with 400 INVALID_DEPLOYMENT; see Errors and limits.