Troubleshooting and FAQ#
Short answers, with a link to the page that explains each in full. For a deployment that waits or fails, start with Troubleshoot a deployment.
Choosing#
llama.cpp or vLLM?#
llama.cpp for quantized GGUF models, small GPUs, CPUs, Macs, AMD Radeon cards vLLM does not support, and a few users at once. vLLM for full or FP8 checkpoints on large NVIDIA or AMD Instinct GPUs and many concurrent users. The engine follows the variant you pick in the library. See Choosing the engine.
Which quantization?#
The largest that fits your GPUs at the context and parallel requests you
need. Q4_K_M, the library's default, loses little; Q5_K_M and Q8_0
are closer to the original. The library shows each one's memory and fit.
See Fit and quantization.
Will this model fit on my machine?#
Open the model's family in Eos → Models → Library and pick the machine under Runs on: each variant says fits on a GPU, needs N GPUs, CPU only or doesn't fit, and why. See Fit and quantization.
Can a replica use GPUs of two machines?#
No. A replica uses 1, 2, 4 or 8 whole GPUs of one machine. Several replicas can run on several machines behind one endpoint.
Can two deployments share one GPU?#
No: a replica holds whole GPUs. Small models can share a machine's CPUs: CPU replicas use the machine's idle cores and several fit on one machine.
Deployments#
My deployment is Pending. Why?#
Press Why? next to its state: it lists every machine and what keeps it out, and what you can do. Usually the GPUs are in use, or none has enough memory for the plan, or the machine has no data location. See Troubleshoot a deployment.
The first start takes long.#
The first replica on a machine downloads the weights (the fill run's log shows the bytes), then the engine loads them. Later starts on that machine skip the download. Pull the model ahead of time to start quickly (Pull onto a machine).
What does scale to zero cost me?#
Nothing while idle: no replica runs and no GPU is held. The first request
afterwards waits while a replica starts — seconds to minutes depending on
the model — and the gateway holds it up to 2 minutes, then answers 503
with Retry-After. Keep min: 1 for interactive use.
Can I change a deployment without downtime?#
Yes, with two or more replicas: changes replace replicas one at a time, and an old one stops only once its replacement serves.
Can I pass my own llama.cpp or vLLM flags?#
Yes, from an allow-list of tuning, sampling and parser flags
(engine_args). Flags that read files, reach the network, set keys or open
the server further are refused. See
Engine arguments.
Can I load a LoRA adapter at run time?#
No. Merge the adapter into the model and serve the merged model; see Serve a fine-tuned model.
Can I serve a model I trained or converted myself?#
Yes: register it from the drive that holds it (source.drive, with
format), or push it to a private Hugging Face repository and add it from
there. See Add a model → From a drive.
Calling deployments#
Where is my endpoint?#
On the deployment's page, under Gateway → Base URL, with curl and
Python examples. If it says no gateway serves the cluster, run the agent's
edge part on a machine your clients can reach, or turn on the hosted
gateway. See Gateways.
Does it work with my OpenAI client or framework?#
Yes, if it lets you set the base URL: the OpenAI SDKs, LangChain,
LlamaIndex, Continue, Aider and others. Put the deployment's name in
model.
Is there an endpoint for speech to text?#
The gateways serve /v1/models, /v1/chat/completions, /v1/completions
and /v1/embeddings. Transcribe in the
Playground, or call a Whisper deployment
from a run of the workspace at its in-cluster endpoint, where vLLM serves
/v1/audio/transcriptions.
How do I use TLS?#
Put your own TLS in front of the gateway (a load balancer, or Caddy or
nginx on the machine) and set ASTRAEUS_GATEWAY_URL. See
An OpenAI-compatible endpoint for your app.
Are there rate limits?#
The gateway on your machines sets none: a replica serves parallel
requests at once, and more wait. The hosted gateway limits calls with
unknown keys (30 per minute per address), bodies (1 MB) and time (90 s).
Why does streaming fail with stream_unsupported?#
You are calling the hosted gateway, which does not stream. Call the gateway
on your machines, or send stream: false.
Security and data#
Do prompts or answers reach Astralyx?#
Not through the gateway on your machines: requests and answers go from your application to your machines only. They pass through Astralyx's console — without being kept — in the Playground and through the hosted gateway, which is off unless a workspace admin turns it on. See What stays on your machines.
Where are the weights?#
On the machines that downloaded them, under the data location their owner chose — or in your own drive, for a model registered from one.
What does Astralyx keep about my keys?#
Only their SHA-256 and first characters. A key is shown once, when it is created.
How are gated models' tokens handled?#
You store your Hugging Face token in your own secret store and add a credential that refers to it. The machine that downloads fetches the token itself; Astralyx keeps only the credential's name.