An OpenAI-compatible endpoint for your app#
You give an application https://llm.example.com/v1: an OpenAI-compatible
endpoint, served by the gateway on one of your machines behind TLS, with
an API key that may call one deployment. Your application keeps using the
OpenAI SDK; only the base URL and the key change. Prompts and answers never
leave your machines.
What you need:
- A deployment that is
Ready, herechat(Serve an open chat model on one GPU machine). - A Linux machine of the cluster that your application can reach — it can
be the GPU machine itself — and a DNS name for it (
llm.example.com). - Root on that machine, for the agent's edge part and a TLS proxy.
- The variables in How the recipes are written.
1. Run the gateway on the machine#
The gateway is part of the agent's edge part. Install it on the machine with the install command the console gives under Machines → Add machine, adding the edge part to the agent's parts:
$ echo '<token>' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh \
| sudo sh -s -- --apiserver <the address shown in the console> --token-file ./astraeus-token \
--agents drives,credentials,data,edge
On a machine already in the cluster, run the installer again with
--agents drives,credentials,data,edge
(Installer,
Update, drain and remove). Envoy is
not needed for the gateway. Then:
$ systemctl is-active astraeus-agent-edge
active
$ curl -sS http://127.0.0.1:8800/v1/models
{"error":{"message":"You didn't provide an API key: send it as `Authorization: Bearer ak-…`.","type":"invalid_request_error","param":null,"code":"missing_api_key"}}
The gateway answers on port 8800. Within a minute the deployment's page lists it under Gateway → Base URL.
2. Put TLS in front of it#
The gateway speaks plain HTTP. Terminate TLS in front of it — here with Caddy on the same machine, which gets a certificate for you — and keep 8800 off the network.
-
Make the gateway listen on loopback only, and tell it the address clients use. In
/etc/astraeus/agent.env: -
Proxy HTTPS to it:
flush_interval -1passes streamed answers through as they come. -
Open TCP 443 to your application's network in the machine's firewall.
The deployment's page and API keys → Gateways now show
https://llm.example.com.
Several gateways
Every machine with the edge part serves the gateway. For redundancy, run it on two machines and put both behind your load balancer; each reaches every replica.
3. Make a key for the application#
One key per application, limited to what it calls, expiring when you rotate keys:
Eos → API keys → New key: Name support-app, May call
Only these: chat, Expires after (days) 90. Press Make the
key and store it in the application's secret store.
4. Call it from the application#
import os
from openai import OpenAI, APIStatusError
client = OpenAI(
base_url="https://llm.example.com/v1",
api_key=os.environ["ASTRAEUS_API_KEY"],
max_retries=5, # 502 and 503 (a replica starting) are retried, honouring Retry-After
timeout=120,
)
def answer(question: str) -> str:
try:
r = client.chat.completions.create(
model="chat",
messages=[
{"role": "system", "content": "You are the support assistant of Acme. Be brief."},
{"role": "user", "content": question},
],
temperature=0.3,
max_tokens=400,
)
return r.choices[0].message.content
except APIStatusError as e:
# e.status_code and e.body["error"]["code"]: invalid_api_key, model_not_allowed, …
raise RuntimeError(f"LLM refused: {e.status_code} {e.body}") from e
print(answer("How do I reset my password?"))
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://llm.example.com/v1",
apiKey: process.env.ASTRAEUS_API_KEY,
maxRetries: 5,
timeout: 120_000,
});
const stream = await client.chat.completions.create({
model: "chat",
messages: [{ role: "user", content: "How do I reset my password?" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");
$ curl -sS https://llm.example.com/v1/chat/completions \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "chat", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'
data: {"choices":[{"delta":{"content":"Hello"},"index":0}],…}
…
data: [DONE]
Everything the engine supports passes through: tools, response_format
(JSON mode and JSON Schema), logprobs, stop, seed. Libraries built on
the OpenAI API — LangChain's ChatOpenAI, LlamaIndex's OpenAILike — work
the same way with base_url and api_key.
5. Check that it worked#
curl -sS https://llm.example.com/v1/models -H "Authorization: Bearer $ASTRAEUS_API_KEY"listschatonly.- A call with another deployment's name answers
403 model_not_allowed. - On Eos → API keys,
support-app's Last used shows the minute of your last call. - Organisation → Usage → Tokens by model counts the tokens.
Handle errors#
Status, code |
What the application should do |
|---|---|
503 model_unavailable (with Retry-After: 5) |
Retry: a replica is starting (the deployment was at zero) or none is serving. The OpenAI SDKs retry it. |
502 upstream_error |
Retry once; if it persists, a replica is failing. |
401 invalid_api_key, expired_api_key |
Do not retry; rotate the key. |
403 model_not_allowed, 404 model_not_found |
Do not retry; fix model or the key's scope. |
413 request_too_large |
The body is over 16 MiB; send less. |
Clean up#
Revoke the key (Eos → API keys → Revoke). To stop serving the gateway
on the machine, set ASTRAEUS_GATEWAY_ADDRESS= (empty) and restart
astraeus-agent-edge.