Skip to content

OpenAI-compatible API#

Eos gateways speak the OpenAI API. Any OpenAI client library works: point its base URL at a gateway, give it an API key, and put the deployment's name in model. This page lists the routes, how requests are routed and checked, streaming, limits and errors.

Base URL#

Gateway Base URL
On your machines http://<machine>:8800/v1, or the address you put in front of it (shown on the deployment's page and under API keys → Gateways)
Hosted (when a workspace admin turned it on) https://console.astralyx.cloud/inference/v1
client.py
import os
from openai import OpenAI

client = OpenAI(base_url="http://10.0.4.12:8800/v1", api_key=os.environ["ASTRAEUS_API_KEY"])

Authentication#

Authorization: Bearer ak-…

A workspace API key. No other header (such as x-api-key) is read.

Routes#

Method and path What
GET /v1/models The deployments the key may call.
POST /v1/chat/completions A chat completion, streamed or not.
POST /v1/completions A text completion, streamed or not.
POST /v1/embeddings Embeddings, for an embedding deployment.

Any other path is answered 404 unknown_url. Speech to text (/v1/audio/transcriptions) is not served by the gateways: transcribe in the Playground, or call a Whisper deployment from a run of the workspace at its endpoint (see Inside the cluster).

The request body is passed to the engine as it is, at the same path (only model is rewritten for a shared deployment). Every field the engine supports works: temperature, top_p, max_tokens, stop, seed, tools, tool_choice, response_format, logprobs, and the engine's own extensions (top_k, min_p, repeat_penalty for llama.cpp; chat_template_kwargs for both). Only the Content-Type and Accept headers are forwarded.

GET /v1/models#

$ curl -sS http://10.0.4.12:8800/v1/models -H "Authorization: Bearer $ASTRAEUS_API_KEY"
{"object":"list","data":[{"id":"chat","object":"model","created":0,"owned_by":"astraeus"},{"id":"ws-5c1e….vision-chat","object":"model","created":0,"owned_by":"ws-5c1e…"}]}

Your workspace's deployments have their own name as id; a deployment another workspace shares with yours has <owner namespace>.<deployment>.

POST /v1/chat/completions#

$ curl -sS http://10.0.4.12:8800/v1/chat/completions \
    -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
    -d '{"model": "chat", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'

Images, audio and video go inline as image_url, input_audio and video_url parts holding data, for a deployment whose model takes them (its capabilities). A model that reasons returns its thoughts as the engine sends them: in the text as <think>…</think>, or apart when vLLM runs with a reasoning parser.

POST /v1/embeddings#

$ curl -sS http://10.0.4.12:8800/v1/embeddings \
    -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
    -d '{"model": "embed", "input": ["The quick brown fox", "A lazy dog"]}'

An embedding deployment answers only this route; a chat deployment does not serve it.

Routing#

  • model is a deployment of the key's workspace, by its name (chat), or a deployment shared with the key's workspace, as <owner namespace>.<deployment>.
  • The request goes to the serving replica with the fewest requests in flight; between equals, in turn.
  • A deployment scaled to zero is woken and the request is held until a replica serves — up to 2 minutes at the gateway on your machines (the edge part's --activator-timeout, ASTRAEUS_ACTIVATOR_TIMEOUT) — then answered 503 model_unavailable with Retry-After: 5.

Streaming#

With "stream": true, server-sent events pass through as the engine sends them, unbuffered. To count tokens, the gateway asks the engine for the streamed answer's usage (stream_options.include_usage): if your client did not ask for it, the final usage chunk is removed before it reaches you; if it did, you get it as usual.

The hosted gateway does not stream: stream: true is refused with 400 stream_unsupported.

Limits#

Limit Gateway on your machines Hosted gateway
Request body 16 MiB 1 MB
Non-streamed answer 64 MiB —
Connecting to a replica 10 s —
Silence in an answer 10 minutes, then it is cut —
Whole call — 90 s, a wake from zero included
Unknown keys — 30 per client address per minute, then 429 rate_limited
Wait for a deployment at zero 2 minutes (configurable) Within the 90 s

The gateways set no limit on requests per second or tokens: a deployment serves parallel requests per replica at once, and more wait at the engine. A client that disconnects frees its replica slot; its tokens so far are counted.

Errors#

Errors have OpenAI's shape, so client libraries raise them as they would OpenAI's:

{"error": {"message": "This API key may not call the model `chat`.", "type": "invalid_request_error", "param": null, "code": "model_not_allowed"}}
Status code Message Cause
400 invalid_request We could not parse the JSON body of your request: …, You must provide a model parameter., model must be a string. The body is not JSON, or has no string model.
400 stream_unsupported Streaming is not available through the hosted gateway: send stream: false, or call your workspace's gateway at the edge, which streams. Hosted gateway only.
401 missing_api_key You didn't provide an API key: send it as Authorization: Bearer ak-…. No key.
401 invalid_api_key Incorrect API key provided. Unknown or revoked key.
401 expired_api_key This API key has expired. Past expires_at.
403 model_not_allowed This API key may not call the model chat. The key is limited to other deployments.
403 model_not_shared The model … is not shared with this key's workspace (or no longer). A shared deployment whose share was revoked, expired or lacks call.
403 hosted_gateway_disabled The hosted gateway is off for this key's workspace… Hosted gateway only.
404 model_not_found The model chat does not exist or you do not have access to it. No such deployment for this key.
404 unknown_url Invalid URL (GET /v1/foo): the gateway serves /v1/models, /v1/chat/completions, /v1/completions and /v1/embeddings. Another path.
413 request_too_large Request bodies are limited to 16 MiB. (hosted: …1 MB…) The body is too large.
429 rate_limited Too many requests with unknown keys; slow down. Hosted gateway only.
502 upstream_error The model chat did not answer; retry. / …stopped answering; retry. A replica was unreachable or cut the answer.
503 model_unavailable The model chat is starting; retry shortly. / No replica of chat is serving; retry shortly. At zero and still starting, or no replica can serve. With Retry-After: 5.
503 unavailable The gateway is unavailable; retry. / The cluster is unavailable; retry. Hosted gateway only.
504 timeout The model chat did not answer within 90 s, the hosted gateway's limit… Hosted gateway only.

Errors from the engine itself (a context too long, an invalid parameter) are passed through with the engine's status and body.

Inside the cluster#

Runs of the same workspace can call a deployment without a key, at its endpoint:

http://deploy-<deployment>.<namespace>.astraeus.local:8000/v1

The deployment's page shows the exact address under Endpoint. This reaches the engine directly, so every route the engine serves works — including /v1/audio/transcriptions for a Whisper deployment — but nothing is counted per key, and a deployment at zero has no address until a request through a gateway wakes it.