Skip to content

Gateways, API keys, sharing and usage#

Programs call a workspace's deployments at one OpenAI-compatible address, a gateway, with an API key of the workspace. Any OpenAI client works: point its base URL at the gateway and put the deployment's name in model. This page explains the two gateways, keys, sharing a deployment with another workspace, and how tokens are counted.

The gateway on your machines#

The gateway runs in the agent's edge part (astraeus-agent-edge) on a machine of your cluster, and listens on port 8800. Requests and answers go from your application to that machine and on to a replica: they never leave your machines. Astralyx sees only token counts and the hashes of your keys.

  • Set it up: install (or reinstall) a machine your clients can reach with the edge part among the agent's parts: --agents drives,credentials,data,edge (Installer). Each such machine serves the gateway; the deployment's page lists them under Gateway.
  • Address: by default http://<machine address>:8800. Change the listen address with ASTRAEUS_GATEWAY_ADDRESS in /etc/astraeus/agent.env (empty turns the gateway off).
  • TLS: the gateway speaks plain HTTP. Put your TLS in front of it — your load balancer, or a proxy on the machine such as Caddy or nginx — and set ASTRAEUS_GATEWAY_URL to the address clients use (https://llm.example.com). The console then shows that address.
  • Firewall: open TCP 8800 (or your proxy's port) to your clients only (Network and firewalls).
  • Streaming: server-sent events pass through as the engine sends them.

The hosted gateway#

For clients that cannot reach your machines, a workspace admin can turn on the hosted gateway (API keys → Gateways → Turn the hosted gateway on). The same keys then work at https://console.astralyx.cloud/inference/v1.

Requests and answers pass through Astralyx

With the hosted gateway, prompts and answers pass through Astralyx's console on their way to your machines. They are not kept, but they do leave your network. It is off by default.

Its limits, compared with the gateway on your machines:

Gateway on your machines Hosted gateway
Address http://<machine>:8800/v1, or yours https://console.astralyx.cloud/inference/v1
Streaming Yes No: stream: true is refused (stream_unsupported)
Request body 16 MiB 1 MB
Time for an answer No limit while the answer keeps coming (cut after 10 minutes of silence) 90 s, a wake from zero included
Routes /v1/models, /v1/chat/completions, /v1/completions, /v1/embeddings The same
Where the data goes Your machines only Through Astralyx

See OpenAI-compatible API for every route, header and error.

API keys#

An API key (ak- followed by 43 characters) belongs to the workspace, not to a person: it keeps working when whoever made it leaves.

  • Shown once. Only its SHA-256 is kept. The list shows its first characters (ak-7Gd2…), who made it, when, and when a gateway last saw it used (to the minute).
  • Scoped. A key may call every deployment of the workspace, now and later, or only those you list. It may also call deployments other workspaces share with yours.
  • Expiring. Optionally, from a date on. An expired key is refused with expired_api_key.
  • Revoked by deleting it. Gateways refuse it from then on.

Editors and admins of the workspace create and revoke keys; viewers see the list. See Manage API keys.

Sharing a deployment#

A workspace can share one of its deployments with another workspace — of the same organisation or another one:

  • Call with an API key: the other workspace's keys may call it through the gateway, naming it <owner namespace>.<deployment> in model.
  • Try in their Playground: their editors can chat with it, embed or transcribe in their own console.

A share lasts until it expires or is revoked. It is checked on every call: after a revocation their calls are refused within half a minute. The receiving workspace sees the deployment's name, model, state and capabilities, never its machines, replicas or engine arguments. Answers come from your machines; their tokens count under your deployment and in their workspace's usage. See Share a deployment.

Tokens and usage#

Every answer through a gateway, and every Playground answer, is counted: prompt tokens, completion tokens and requests, per deployment, key and minute. The counts reach your usage within seconds, under the deployment's model:

  • Organisation → Usage shows tokens in and out next to GPU-hours, per workspace, and a Tokens by model table with requests, prompt and completion tokens and their cost.
  • Calls to a deployment another workspace shares with yours are counted in your usage as shared:<owner namespace>.<model>.
  • A request counts even when it fails; tokens count only for answers the engine reported usage for. A streamed answer's usage is requested from the engine by the gateway; a client that did not ask for it does not see it.

To put a price on tokens, an organisation admin sets prices per million tokens in the cluster's price table: tokens_prompt:<model> and tokens_completion:<model>, or tokens_prompt and tokens_completion for any model. Replicas also hold GPUs, which are priced as GPU-hours like any run; a price table usually prices one or the other. See Usage, cost and budgets and Watch usage.

What stays on your machines#

Where it is
Weights On the machines that downloaded them, under the data location their owner chose
Prompts and answers through the gateway on your machines On your machines only: the gateway machine and the replica
Prompts and answers in the Playground or through the hosted gateway They pass through Astralyx's console, which keeps none of them
Playground conversations In your browser
API keys Shown once; only their SHA-256 is kept
Hugging Face tokens In your secret store, read by the machine that downloads
What Astralyx keeps Names, specifications, states, token counts and usage

See What leaves your machines.