Gateways, API keys, sharing and usage#
Programs call a workspace's deployments at one OpenAI-compatible address, a
gateway, with an API key of the workspace. Any OpenAI client works:
point its base URL at the gateway and put the deployment's name in model.
This page explains the two gateways, keys, sharing a deployment with another
workspace, and how tokens are counted.
The gateway on your machines#
The gateway runs in the agent's edge part (astraeus-agent-edge) on a
machine of your cluster, and listens on port 8800. Requests and answers
go from your application to that machine and on to a replica: they never
leave your machines. Astralyx sees only token counts and the hashes of
your keys.
- Set it up: install (or reinstall) a machine your clients can reach
with the edge part among the agent's parts:
--agents drives,credentials,data,edge(Installer). Each such machine serves the gateway; the deployment's page lists them under Gateway. - Address: by default
http://<machine address>:8800. Change the listen address withASTRAEUS_GATEWAY_ADDRESSin/etc/astraeus/agent.env(empty turns the gateway off). - TLS: the gateway speaks plain HTTP. Put your TLS in front of it — your
load balancer, or a proxy on the machine such as Caddy or nginx — and set
ASTRAEUS_GATEWAY_URLto the address clients use (https://llm.example.com). The console then shows that address. - Firewall: open TCP 8800 (or your proxy's port) to your clients only (Network and firewalls).
- Streaming: server-sent events pass through as the engine sends them.
The hosted gateway#
For clients that cannot reach your machines, a workspace admin can turn on
the hosted gateway (API keys → Gateways → Turn the hosted gateway
on). The same keys then work at
https://console.astralyx.cloud/inference/v1.
Requests and answers pass through Astralyx
With the hosted gateway, prompts and answers pass through Astralyx's console on their way to your machines. They are not kept, but they do leave your network. It is off by default.
Its limits, compared with the gateway on your machines:
| Gateway on your machines | Hosted gateway | |
|---|---|---|
| Address | http://<machine>:8800/v1, or yours |
https://console.astralyx.cloud/inference/v1 |
| Streaming | Yes | No: stream: true is refused (stream_unsupported) |
| Request body | 16 MiB | 1 MB |
| Time for an answer | No limit while the answer keeps coming (cut after 10 minutes of silence) | 90 s, a wake from zero included |
| Routes | /v1/models, /v1/chat/completions, /v1/completions, /v1/embeddings |
The same |
| Where the data goes | Your machines only | Through Astralyx |
See OpenAI-compatible API for every route, header and error.
API keys#
An API key (ak- followed by 43 characters) belongs to the workspace,
not to a person: it keeps working when whoever made it leaves.
- Shown once. Only its SHA-256 is kept. The list shows its first
characters (
ak-7Gd2…), who made it, when, and when a gateway last saw it used (to the minute). - Scoped. A key may call every deployment of the workspace, now and later, or only those you list. It may also call deployments other workspaces share with yours.
- Expiring. Optionally, from a date on. An expired key is refused with
expired_api_key. - Revoked by deleting it. Gateways refuse it from then on.
Editors and admins of the workspace create and revoke keys; viewers see the list. See Manage API keys.
Sharing a deployment#
A workspace can share one of its deployments with another workspace — of the same organisation or another one:
- Call with an API key: the other workspace's keys may call it through
the gateway, naming it
<owner namespace>.<deployment>inmodel. - Try in their Playground: their editors can chat with it, embed or transcribe in their own console.
A share lasts until it expires or is revoked. It is checked on every call: after a revocation their calls are refused within half a minute. The receiving workspace sees the deployment's name, model, state and capabilities, never its machines, replicas or engine arguments. Answers come from your machines; their tokens count under your deployment and in their workspace's usage. See Share a deployment.
Tokens and usage#
Every answer through a gateway, and every Playground answer, is counted: prompt tokens, completion tokens and requests, per deployment, key and minute. The counts reach your usage within seconds, under the deployment's model:
- Organisation → Usage shows tokens in and out next to GPU-hours, per workspace, and a Tokens by model table with requests, prompt and completion tokens and their cost.
- Calls to a deployment another workspace shares with yours are counted in
your usage as
shared:<owner namespace>.<model>. - A request counts even when it fails; tokens count only for answers the engine reported usage for. A streamed answer's usage is requested from the engine by the gateway; a client that did not ask for it does not see it.
To put a price on tokens, an organisation admin sets prices per million
tokens in the cluster's price table: tokens_prompt:<model> and
tokens_completion:<model>, or tokens_prompt and tokens_completion for
any model. Replicas also hold GPUs, which are priced as GPU-hours like any
run; a price table usually prices one or the other. See
Usage, cost and budgets and
Watch usage.
What stays on your machines#
| Where it is | |
|---|---|
| Weights | On the machines that downloaded them, under the data location their owner chose |
| Prompts and answers through the gateway on your machines | On your machines only: the gateway machine and the replica |
| Prompts and answers in the Playground or through the hosted gateway | They pass through Astralyx's console, which keeps none of them |
| Playground conversations | In your browser |
| API keys | Shown once; only their SHA-256 is kept |
| Hugging Face tokens | In your secret store, read by the machine that downloads |
| What Astralyx keeps | Names, specifications, states, token counts and usage |
See What leaves your machines.