OpenAI-compatible API#
Eos gateways speak the OpenAI API. Any OpenAI client library works: point
its base URL at a gateway, give it an API key, and put the deployment's name
in model. This page lists the routes, how requests are routed and
checked, streaming, limits and errors.
Base URL#
| Gateway | Base URL |
|---|---|
| On your machines | http://<machine>:8800/v1, or the address you put in front of it (shown on the deployment's page and under API keys → Gateways) |
| Hosted (when a workspace admin turned it on) | https://console.astralyx.cloud/inference/v1 |
import os
from openai import OpenAI
client = OpenAI(base_url="http://10.0.4.12:8800/v1", api_key=os.environ["ASTRAEUS_API_KEY"])
Authentication#
A workspace API key. No other header (such as
x-api-key) is read.
Routes#
| Method and path | What |
|---|---|
GET /v1/models |
The deployments the key may call. |
POST /v1/chat/completions |
A chat completion, streamed or not. |
POST /v1/completions |
A text completion, streamed or not. |
POST /v1/embeddings |
Embeddings, for an embedding deployment. |
Any other path is answered 404 unknown_url. Speech to text
(/v1/audio/transcriptions) is not served by the gateways: transcribe in
the Playground, or call a Whisper
deployment from a run of the workspace at its endpoint (see
Inside the cluster).
The request body is passed to the engine as it is, at the same path
(only model is rewritten for a shared deployment). Every field the engine
supports works: temperature, top_p, max_tokens, stop, seed,
tools, tool_choice, response_format, logprobs, and the engine's own
extensions (top_k, min_p, repeat_penalty for llama.cpp;
chat_template_kwargs for both). Only the Content-Type and Accept
headers are forwarded.
GET /v1/models#
$ curl -sS http://10.0.4.12:8800/v1/models -H "Authorization: Bearer $ASTRAEUS_API_KEY"
{"object":"list","data":[{"id":"chat","object":"model","created":0,"owned_by":"astraeus"},{"id":"ws-5c1e….vision-chat","object":"model","created":0,"owned_by":"ws-5c1e…"}]}
Your workspace's deployments have their own name as id; a deployment
another workspace shares with yours has <owner namespace>.<deployment>.
POST /v1/chat/completions#
$ curl -sS http://10.0.4.12:8800/v1/chat/completions \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "chat", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'
Images, audio and video go inline as image_url, input_audio and
video_url parts holding data, for a deployment whose model takes them
(its capabilities). A model that reasons returns its thoughts as the
engine sends them: in the text as <think>…</think>, or apart when vLLM
runs with a reasoning parser.
POST /v1/embeddings#
$ curl -sS http://10.0.4.12:8800/v1/embeddings \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "embed", "input": ["The quick brown fox", "A lazy dog"]}'
An embedding deployment answers only this route; a chat deployment does not serve it.
Routing#
modelis a deployment of the key's workspace, by its name (chat), or a deployment shared with the key's workspace, as<owner namespace>.<deployment>.- The request goes to the serving replica with the fewest requests in flight; between equals, in turn.
- A deployment scaled to zero is woken and the request is held until a
replica serves — up to 2 minutes at the gateway on your machines (the
edge part's
--activator-timeout,ASTRAEUS_ACTIVATOR_TIMEOUT) — then answered503 model_unavailablewithRetry-After: 5.
Streaming#
With "stream": true, server-sent events pass through as the engine sends
them, unbuffered. To count tokens, the gateway asks the engine for the
streamed answer's usage (stream_options.include_usage): if your client
did not ask for it, the final usage chunk is removed before it reaches you;
if it did, you get it as usual.
The hosted gateway does not stream: stream: true is refused with
400 stream_unsupported.
Limits#
| Limit | Gateway on your machines | Hosted gateway |
|---|---|---|
| Request body | 16 MiB | 1 MB |
| Non-streamed answer | 64 MiB | — |
| Connecting to a replica | 10 s | — |
| Silence in an answer | 10 minutes, then it is cut | — |
| Whole call | — | 90 s, a wake from zero included |
| Unknown keys | — | 30 per client address per minute, then 429 rate_limited |
| Wait for a deployment at zero | 2 minutes (configurable) | Within the 90 s |
The gateways set no limit on requests per second or tokens: a deployment
serves parallel requests per replica at once, and more wait at the
engine. A client that disconnects frees its replica slot; its tokens so far
are counted.
Errors#
Errors have OpenAI's shape, so client libraries raise them as they would OpenAI's:
{"error": {"message": "This API key may not call the model `chat`.", "type": "invalid_request_error", "param": null, "code": "model_not_allowed"}}
| Status | code |
Message | Cause |
|---|---|---|---|
| 400 | invalid_request |
We could not parse the JSON body of your request: …, You must provide a model parameter., model must be a string. |
The body is not JSON, or has no string model. |
| 400 | stream_unsupported |
Streaming is not available through the hosted gateway: send stream: false, or call your workspace's gateway at the edge, which streams. |
Hosted gateway only. |
| 401 | missing_api_key |
You didn't provide an API key: send it as Authorization: Bearer ak-…. |
No key. |
| 401 | invalid_api_key |
Incorrect API key provided. | Unknown or revoked key. |
| 401 | expired_api_key |
This API key has expired. | Past expires_at. |
| 403 | model_not_allowed |
This API key may not call the model chat. |
The key is limited to other deployments. |
| 403 | model_not_shared |
The model … is not shared with this key's workspace (or no longer). |
A shared deployment whose share was revoked, expired or lacks call. |
| 403 | hosted_gateway_disabled |
The hosted gateway is off for this key's workspace… | Hosted gateway only. |
| 404 | model_not_found |
The model chat does not exist or you do not have access to it. |
No such deployment for this key. |
| 404 | unknown_url |
Invalid URL (GET /v1/foo): the gateway serves /v1/models, /v1/chat/completions, /v1/completions and /v1/embeddings. | Another path. |
| 413 | request_too_large |
Request bodies are limited to 16 MiB. (hosted: …1 MB…) | The body is too large. |
| 429 | rate_limited |
Too many requests with unknown keys; slow down. | Hosted gateway only. |
| 502 | upstream_error |
The model chat did not answer; retry. / …stopped answering; retry. |
A replica was unreachable or cut the answer. |
| 503 | model_unavailable |
The model chat is starting; retry shortly. / No replica of chat is serving; retry shortly. |
At zero and still starting, or no replica can serve. With Retry-After: 5. |
| 503 | unavailable |
The gateway is unavailable; retry. / The cluster is unavailable; retry. | Hosted gateway only. |
| 504 | timeout |
The model chat did not answer within 90 s, the hosted gateway's limit… |
Hosted gateway only. |
Errors from the engine itself (a context too long, an invalid parameter) are passed through with the engine's status and body.
Inside the cluster#
Runs of the same workspace can call a deployment without a key, at its endpoint:
The deployment's page shows the exact address under Endpoint. This
reaches the engine directly, so every route the engine serves works —
including /v1/audio/transcriptions for a Whisper deployment — but nothing
is counted per key, and a deployment at zero has no address until a request
through a gateway wakes it.