Serve models for the whole company#
One platform team owns the GPUs and the model; every other team calls it. This recipe deploys a coding model with several replicas behind the gateway on your own machines, shares the deployment with three teams' workspaces, and gives each its own API key — so you can watch who uses what, and revoke one team without touching the others.
This recipe is about operating a shared deployment at company scale. For scaling and quantizing the model itself, see Multiple replicas behind one endpoint and Fit a big model with quantization; for the gateway's own reference, see Gateways, API keys, sharing and usage.
Before you begin#
- The editor or admin role in the platform team's workspace
(
platform), which owns the machines and the deployment. - A model already in that workspace, or one from the library (Add a model).
- A machine running the agent's edge part, reachable from every team's network — this is what serves the gateway.
- For the API examples,
ASTRALYX_APIandASTRALYX_TOKENas in Add a model.
1. Make sure a gateway is reachable#
The gateway runs in the agent's edge part and listens on port 8800; requests and answers never leave your machines. If none of your machines run it yet, add it when you install one, or reinstall one with the edge part added to its others:
$ echo '9c4e…' > ./astraeus-token && curl -fsSL https://console.astralyx.cloud/api/v1/install.sh | sudo sh -s -- --control-plane https://connect.astralyx.cloud --token-file ./astraeus-token --releases https://console.astralyx.cloud/releases --org 'Acme Research' --org-id 0192f0c4-7d1e-7a51-9c33-5e8b2a4f6d10 --agents drives,credentials,data,edge && rm -f ./astraeus-token
Put your own TLS in front of it (your load balancer, or Caddy or nginx on
the machine) and set ASTRAEUS_GATEWAY_URL in
/etc/astraeus/agent.env to the address teams should use
(https://llm.acme.internal); the console then shows that address. Open
the port to every team's network, not the whole company's. See
The gateway on your machines.
2. Deploy with several replicas and scale to zero#
$ astra eos deploy qwen2-5-coder-7b-q4-k-m --name coder --min 1 --max 4
coder deployed: `astra eos deployments` shows when it is Ready
For settings the CLI has no flag for — the GPU model, parallel requests per replica, engine arguments — use the console's New deployment form or the API:
{
"metadata": {"name": "coder"},
"spec": {
"model": "qwen2-5-coder-7b-q4-k-m",
"replicas": {"min": 1, "max": 4, "idle_minutes": 30},
"context_length": 16384
}
}
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d @deployment.json
With min: 1, one replica is always up; max: 4 adds replicas as requests
queue. Set min: 0 instead if the whole company's usage is bursty enough
that the first request of the day waiting for a cold start is acceptable:
the deployment scales to none after idle_minutes and the next request
wakes it. See Scale.
3. Share it with each team's workspace#
A share lets another workspace call your deployment through the gateway, or try it in their own Playground, without its own GPUs.
On the deployment's page, under Sharing, Share. Choose A workspace, give its organisation and workspace, tick call with an API key and try in their Playground, and Share. Repeat for each team.
$ curl -fsS "$ASTRALYX_API/eos/share-targets?org=<org>&workspace=vision" -H "Authorization: Bearer $ASTRALYX_TOKEN"
{"org":"<org>","workspace":"vision","namespace":"ws-3f9c2a1b7d4e"}
$ curl -fsS -X POST "$ASTRALYX_API/eos/shares" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d '{
"metadata": {"name": "for-vision"},
"spec": {"deployment": "coder", "to": {"org": "<org>", "workspace": "vision", "namespace": "ws-3f9c2a1b7d4e"}, "access": ["call", "playground"]}
}'
Repeat for research and eval with their own namespaces and share
names (for-research, for-eval).
Answers from your machines; their tokens count under your deployment, and
in theirs as shared:<platform namespace>.coder. They never see your
machines, replicas or engine arguments. See
Share a deployment.
4. Give each team its own key#
Either make a key for them, scoped to this deployment only, or have them make their own.
On the share's row, Create a key for them; hand it over, and revoke it on API keys when they no longer need it.
In their workspace: Eos → API keys → New key, tick Also
deployments shared with this workspace, tick coder. With the API,
they list it as spec.shared_deployments: ["<platform namespace>.coder"]:
Either way, each team's key is revocable on its own: cut off Eval without touching Vision or Research.
5. Call it#
With a key of their own, each team puts <platform namespace>.coder in
model and calls your cluster's gateway:
$ export EOS_URL=https://llm.acme.internal/v1
$ export ASTRALYX_API_KEY=ak-…
$ curl -sS $EOS_URL/chat/completions \
-H "Authorization: Bearer $ASTRALYX_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "ws-3f9c2a1b7d4e.coder", "messages": [{"role": "user", "content": "Write a Python function that reverses a string."}]}'
A key you made for them, scoped to the deployment directly, uses the plain
name coder instead. Any OpenAI client works the same way: point its base
URL at your gateway.
6. Watch who uses what#
Organisation → Usage breaks tokens and cost down by workspace;
Eos → API keys shows each key's Last used. In the owning
workspace's usage, calls from a sharing workspace show under the
deployment's own model name; in each receiving workspace's usage, they
show as shared:<platform namespace>.coder. See
Tokens and usage and
Usage, cost and budgets.
Clean up#
Revoke a team's share — their calls are refused within half a minute — or delete the deployment, which deletes every share with it:
$ curl -fsS -X POST "$ASTRALYX_API/eos/shares/for-eval/revoke" -H "Authorization: Bearer $ASTRALYX_TOKEN"
What you get#
- One set of GPUs serving the whole company, scaling with demand instead of being carved up ahead of time.
- Every team with its own revocable key and its own line in usage, calling a gateway that never sends a prompt outside your network.
- A change to the deployment (more replicas, a bigger context, a newer model) that reaches every team at once.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
403 model_not_shared |
The share was revoked, expired, or does not grant call. |
Share again with call with an API key. |
404 model_not_found |
The caller used the plain deployment name instead of <namespace>.<deployment>, or their key does not list it. |
Use the namespaced id; add the deployment to the key's shared_deployments. |
403 model_not_allowed |
The key may not call this deployment at all. | Use a key scoped to it, or one for every deployment. |
stream_unsupported |
The call went to the hosted gateway, which does not stream. | Call the gateway on your own machines instead. |
| A team's first request of the day is slow | The deployment scaled to zero and is waking up. | Expected with min: 0; set min: 1 for an always-on replica if that cost is worth avoiding. |
| A revoked team can still call for up to half a minute | Shares and keys are checked on every call, with a short cache. | Expected; wait, or revoke earlier for a hard deadline. |