Skip to content

Troubleshoot a deployment#

Every deployment and each of its replicas carries a reason, in words. Start there. This page explains how to read it, what the common reasons mean, and what to do.

Why is it not running?#

On Eos → Deployments, a deployment that waits shows Why? next to its state; on its page the same link sits beside the state (Why are replicas waiting? when it is Ready but some replicas wait).

The dialog, Why … is not running, shows:

  • with several waiting replicas, N of M replicas waiting and a chip per replica;
  • a headline and the advice for it, for example Next to run: starts when running work ends — It needs 8 GPUs of ≥ 8 GB, 32 CPU cores, 42 GiB of memory. The closest is gpu-h100-01: 4 of 8 GPUs free; it needs 8. — with Technical details;
  • Machines: every machine and What keeps it out, the closest marked, and what holds its GPUs;
  • What you can do: only the steps that apply.

The Why? dialog for a deployment whose replicas wait

$ astra inference deployments
DEPLOYMENT   MODEL                 KIND   STATE     REPLICAS  GATEWAY
llama-wide   llama3-1-70b-q4-k-m   chat   Ready     1/3-3     http://10.0.4.12:8800/v1
$ astra astraeus workers deploy-llama-wide

A deployment that is not Ready prints its reason on the next line. The waiting replicas' workers (deploy-<name>-<index>-0) carry the precise reason in their REASON column.

$ curl -fsS "$API/deployments/llama-wide" -H "Authorization: Bearer $ASTRAEUS_TOKEN" \
    | jq '{state: .status.state, reason: .status.reason, replicas: [.members[] | {run, machine, state, reason, ready}]}'

Reasons and fixes#

Pending#

Reason Cause Fix
Waiting for a machine with a GPU with at least 15 GB (or N GPUs with at least X GB each, on one machine, or N CPU cores) No machine your workspace may use has that much free. See Why?: wait for running work to end, free GPUs, or lower what each replica needs — a smaller quantization, a shorter context, fewer requests at once, fewer GPUs per replica.
Choose where to keep data on gpu-01 The machine has no data location, so the weights cannot go there. An organisation admin chooses one on the machine's page.
Waiting for the workspace's quota Your workspace's GPU quota on the cluster is used. Stop other work, or ask an admin for more quota (Quotas).
Next to run… It is next in the queue, behind running work. Wait, or raise the deployment's priority if your workspace may.
A machine listed as has no GPUs, or the GPU maker not accepted gpu_vendor excludes those GPUs, or the engine does not run on them (vLLM on a Mac, or on an AMD GPU its ROCm build does not support). Use llama.cpp and a GGUF variant, or another machine.

Starting#

Reason Meaning
Preparing on gpu-01: the image and the model's weights The engine's image is pulled and the weights are downloaded onto the machine. The first start of a large model takes as long as its download; the fill run's log shows the progress.
Waiting for drive model-… to be filled on gpu-01 The download runs first; the replica starts once the copy is whole.
Starting the engine on gpu-01 The container starts.
Loading the model on gpu-01 The engine loads the weights onto the GPU; it serves once its health check passes.

Failed#

Reason Fix
Model … does not exist The model was deleted. Add it again, or change the deployment's model.
Model …: Pulling onto gpu-01 failed: … The download failed. Read the fill run's log (a gated model without a valid credential, a network error, a checksum mismatch); then Pull again.
Replica deploy-…-0 keeps exiting: … The engine exits: 3 restarts without serving. Open the replica's Log. Usual causes: not enough GPU memory (raise gpus, lower context_length or parallel), an engine argument the engine rejects, a checkpoint the engine cannot load.
llama.cpp serves GGUF; model … is a safetensors checkpoint: serve it with vllm Set engine to auto or vllm, or use a GGUF variant.
spec.context_length 200000: model … takes at most 131072 tokens Lower context_length.

ScaledToZero#

No replica runs until the first request, or Idle for 15 minutes: no replica runs; the next request starts one. This is expected with min: 0. Send a request (or Try it, or the Playground) to start a replica. If you need answers at once, set min to 1.

The gateway answers an error#

Error Cause Fix
503 model_unavailable, The model chat is starting; retry shortly. with Retry-After: 5 The deployment was at zero, and its replica did not serve within the gateway's wait (2 minutes by default). Retry after a few seconds. For large models, keep min: 1.
503 model_unavailable, No replica of chat is serving; retry shortly. No replica is ready and none can be woken (it is Pending or Failed). See the deployment's reason.
502 upstream_error, did not answer / stopped answering A replica was unreachable, or cut the answer (restart, out of memory, 10 minutes of silence). Retry; read the replica's log.
404 model_not_found No deployment of that name in the key's workspace. Use the deployment's name (see Model name on its page).
413 request_too_large The body is over 16 MiB (1 MB through the hosted gateway). Send less; for images, resize them.
400 stream_unsupported stream: true through the hosted gateway. Use the gateway on your machines, or stream: false.
504 timeout The hosted gateway gives a call 90 s. Ask for fewer tokens, or call the gateway on your machines with stream: true.

The keys' errors (401, 403) are in Manage API keys. Every code is in OpenAI-compatible API.

Slow answers#

  • First token slow on every request: long prompts on a CPU replica, or a model larger than its GPU on llama.cpp. Check Each replica on the deployment's page and the fit in the library.
  • Throughput drops with users: llama.cpp serves parallel requests at once (4 by default); more wait. Raise parallel (more cache memory), add replicas (max), or serve the model with vLLM.
  • The first request after a while is slow: the deployment scaled to zero. Keep min: 1, or raise idle_minutes.