Management API
Everything the console does in Eos goes through the same REST API as
Astraeus, at https://console.astralyx.cloud/api/v1, in JSON. This page
lists the Eos endpoints. Authentication, conventions and the error format
are in the Astraeus REST API; calling
deployments is in OpenAI-compatible API.
Base URLs
| Scope |
URL |
| A workspace on a cluster |
https://console.astralyx.cloud/api/v1/orgs/{org}/workspaces/{ws}/clusters/{cluster}/api — written $API below |
| A workspace |
https://console.astralyx.cloud/api/v1/orgs/{org}/workspaces/{ws} — written $CONSOLE below |
| The console |
https://console.astralyx.cloud/api/v1 |
Authenticate with Authorization: Bearer <API token> (create one under
Account → API tokens). Names in paths and bodies are local to your
workspace: chat, not the namespace.
Roles
| Role in the workspace |
May |
| viewer |
List and read models, deployments, API keys and shares. |
| editor, admin |
Everything below: add, pull and delete models; create, change and delete deployments; create and revoke keys and shares; use the Playground. |
| admin |
Also turn the hosted gateway on or off. |
Models
| Method and path |
Description |
POST $API/models |
Create a model: {metadata: {name}, spec} (Model specification). 201. |
GET $API/models |
{items: [...]}: every model with its state, copies, pulls and memory estimate. |
GET $API/models/{name} |
One model. |
POST $API/models/{name}/pull |
Download its weights onto a machine: {"machine": "gpu-01"} (optional). 202 with {machine, run, outcome}; outcome is created, running or done. |
DELETE $API/models/{name} |
Delete it (and, from the library or Hugging Face, its weights on every machine). ?force=true even while workers use them. 204. |
The model library
| Method and path |
Description |
GET /catalog/models |
Every library entry: {version, generated_at, items, aliases}. Each item has id (family:size-tag), display, family, publisher, params_b, format, engine, quantization, context_length, license, gated, size_bytes, min_gpu_memory_gb, cpu_ok, capabilities, default, default_name. |
GET $CONSOLE/library |
Families with a fit for your machines. Query: q, category, publisher, size (le3, 3-9, 9-20, 20-40, 40plus), engine, license (permissive, restricted), hide_gated, machines (comma-separated), fits_only, sort (popular, newest, smallest, name), cluster, context (default 8192). |
GET $CONSOLE/library/{family} |
One family: its sizes and variants, each with ref, format, engine, quantization, size_bytes, need, fit (machine by machine) and added_as. |
POST $CONSOLE/clusters/{cluster}/models |
Add a library entry to the workspace: {"catalog": "llama3.2:3b-q4_k_m", "name"?, "credential"?}. 201 with the model. |
GET /catalog/huggingface?repo=owner/name&revision=main |
A public Hugging Face repository's files at the commit the revision points to: {repo, revision, gated, files: [{path, size_bytes, sha256}]}. |
A fit is {status, gpus, parallel, default_parallel, slow, variant,
summary, machines: [{machine, fits, text, backend}]}, where status is
gpu, gpus, cpu, none or unknown.
Deployments
| Method and path |
Description |
POST $API/deployments |
Create: {metadata: {name}, spec} (Deployment specification). 201. |
GET $API/deployments |
{items: [...]}, each with its state, replicas, endpoint, plan, gateways and capabilities. |
GET $API/deployments/{name} |
One deployment. |
PUT $API/deployments/{name} |
Change: {spec}, the whole specification. Replicas are replaced one at a time. |
DELETE $API/deployments/{name} |
Delete it and its shares; the model stays. 204. |
API keys
| Method and path |
Description |
POST $API/api-keys |
Create: {metadata: {name}, spec: {deployments?, shared_deployments?, expires_at?}}. 201 with key, shown once. |
GET $API/api-keys |
{items: [...]}: metadata, spec, prefix, created_by, created_at, last_used, expired. |
GET $API/api-keys/{name} |
One key, without its secret. |
DELETE $API/api-keys/{name} |
Revoke it. 204. |
PATCH $CONSOLE |
{"hosted_gateway": true} or false (workspace admins). The workspace then shows hosted_gateway_url. |
Shares
| Method and path |
Description |
GET $CONSOLE/share-targets?org=&workspace= |
Look up a workspace to share with: {org, org_name, workspace, workspace_name, namespace}. |
POST $API/deployment-shares |
Share: {metadata: {name}, spec: {deployment, to: {org, workspace, namespace} \| to_namespace, access: ["call", "playground"], expires_at?, note?}}. 201. |
GET $API/deployment-shares |
{items: [...]}; ?deployment=<name> for one deployment's. Each has status.state Active or Revoked. |
GET $API/deployment-shares/{name} |
One share. |
POST $API/deployment-shares/{name}/revoke |
Revoke it. |
DELETE $API/deployment-shares/{name} |
Delete it. 204. |
GET $CONSOLE/shared-deployments |
Deployments shared with this workspace: {items: [{cluster, share, deployment, owner_namespace, owner, access, expires_at, note, state, kind, model, capabilities}], unreachable}. |
The Playground's routes
The Playground uses these routes; scripts may too (editors and admins).
| Method and path |
Description |
POST $API/deployments/{name}/chat |
Start an answer: an OpenAI chat request (messages, sampling fields, tools, response_format, chat_template_kwargs, reasoning_effort). 202 with {id, state}. |
GET $API/deployments/{name}/chat/{id}?after=<offset> |
What was written since offset: {text, offset, done, state, finish_reason, usage, error, timings}. Read until done. |
DELETE $API/deployments/{name}/chat/{id} |
Stop it. |
POST $API/deployments/{name}/embeddings |
{"input": [...]}, at most 64 texts: OpenAI's embeddings answer. |
POST $API/deployments/{name}/transcriptions |
Speech to text, for a Whisper deployment. |
POST $CONSOLE/shared-deployments/{cluster}/{deployment}/chat, …/chat/{id}, …/embeddings, …/transcriptions |
The same, for a deployment shared with this workspace for the Playground. |
model and stream are the deployment's to decide. A request to a
deployment at zero wakes it; the answer's state is starting until a
replica serves.
Runs that use a model
A run's specification takes model (see
Batch runs that read a model):
| Field |
Type |
Default |
Description |
spec.model.name |
string |
required |
A model of the run's workspace. |
spec.model.mount_path |
string |
/model |
Where its weights are mounted, read-only. |
spec.model.engine |
bool |
false |
Run the model's engine beside the run's workers. |
Errors
| Code |
Status |
Meaning |
INVALID_MODEL |
400 |
The model specification is invalid; the message names the field. |
MODEL_NOT_FOUND |
404 |
No such model (also when a deployment names one). |
MODEL_ALREADY_EXISTS |
409 |
A model of that name exists. |
MODEL_IN_USE |
409 |
Live workers use its weights; delete with ?force=true. |
MODEL_CONFLICT |
409 |
A drive of yours has the name the model's drive would take. |
NO_DATA_LOCATION |
409 |
The machine chosen for a pull has no data location. |
NO_ROOM |
409 |
No machine has room for the weights, or none may hold them. |
NODE_NOT_FOUND |
404 |
No such machine for this workspace. |
CATALOG_ENTRY_NOT_FOUND |
404 |
No library entry by that reference. |
CREDENTIAL_REQUIRED |
400 |
A gated library entry needs a credential. |
FAMILY_NOT_FOUND |
404 |
No library family by that name. |
INVALID_REPO |
400 |
Not owner/name, or an invalid revision. |
REPO_GATED |
403 |
The repository is private or gated: add it with its files listed by hand, and a credential. |
REPO_NOT_FOUND |
404 |
No such repository or revision on Hugging Face. |
HUB_UNREACHABLE |
503 |
Hugging Face did not answer. |
INVALID_DEPLOYMENT |
400 |
The deployment specification is invalid, or does not fit its model. |
DEPLOYMENT_NOT_FOUND |
404 |
No such deployment. |
DEPLOYMENT_ALREADY_EXISTS |
409 |
A deployment of that name exists. |
INVALID_API_KEY |
400 |
The key's specification is invalid. |
API_KEY_NOT_FOUND |
404 |
No such key. |
API_KEY_ALREADY_EXISTS |
409 |
A key of that name exists. |
INVALID_DEPLOYMENT_SHARE |
400 |
The share's specification is invalid. |
DEPLOYMENT_SHARE_NOT_FOUND, DEPLOYMENT_SHARE_ALREADY_EXISTS |
404, 409 |
— |
SHARE_TO_SELF |
400 |
Sharing with the deployment's own workspace. |
SHARED_DEPLOYMENT_NOT_FOUND |
404 |
No such deployment is shared with your workspace (or not for this). |
NO_READY_BACKEND |
503 |
The deployment has no serving replica and could not be woken in time; retry. |
WORKSPACE_ROLE |
403 |
The route is for the workspace's editors. |