Skip to content

Management API#

Everything the console does in Eos goes through the same REST API as Astraeus, at https://console.astralyx.cloud/api/v1, in JSON. This page lists the Eos endpoints. Authentication, conventions and the error format are in the Astraeus REST API; calling deployments is in OpenAI-compatible API.

Base URLs#

Scope URL
A workspace on a cluster https://console.astralyx.cloud/api/v1/orgs/{org}/workspaces/{ws}/clusters/{cluster}/api — written $API below
A workspace https://console.astralyx.cloud/api/v1/orgs/{org}/workspaces/{ws} — written $CONSOLE below
The console https://console.astralyx.cloud/api/v1

Authenticate with Authorization: Bearer <API token> (create one under Account → API tokens). Names in paths and bodies are local to your workspace: chat, not the namespace.

Roles#

Role in the workspace May
viewer List and read models, deployments, API keys and shares.
editor, admin Everything below: add, pull and delete models; create, change and delete deployments; create and revoke keys and shares; use the Playground.
admin Also turn the hosted gateway on or off.

Models#

Method and path Description
POST $API/models Create a model: {metadata: {name}, spec} (Model specification). 201.
GET $API/models {items: [...]}: every model with its state, copies, pulls and memory estimate.
GET $API/models/{name} One model.
POST $API/models/{name}/pull Download its weights onto a machine: {"machine": "gpu-01"} (optional). 202 with {machine, run, outcome}; outcome is created, running or done.
DELETE $API/models/{name} Delete it (and, from the library or Hugging Face, its weights on every machine). ?force=true even while workers use them. 204.

The model library#

Method and path Description
GET /catalog/models Every library entry: {version, generated_at, items, aliases}. Each item has id (family:size-tag), display, family, publisher, params_b, format, engine, quantization, context_length, license, gated, size_bytes, min_gpu_memory_gb, cpu_ok, capabilities, default, default_name.
GET $CONSOLE/library Families with a fit for your machines. Query: q, category, publisher, size (le3, 3-9, 9-20, 20-40, 40plus), engine, license (permissive, restricted), hide_gated, machines (comma-separated), fits_only, sort (popular, newest, smallest, name), cluster, context (default 8192).
GET $CONSOLE/library/{family} One family: its sizes and variants, each with ref, format, engine, quantization, size_bytes, need, fit (machine by machine) and added_as.
POST $CONSOLE/clusters/{cluster}/models Add a library entry to the workspace: {"catalog": "llama3.2:3b-q4_k_m", "name"?, "credential"?}. 201 with the model.
GET /catalog/huggingface?repo=owner/name&revision=main A public Hugging Face repository's files at the commit the revision points to: {repo, revision, gated, files: [{path, size_bytes, sha256}]}.

A fit is {status, gpus, parallel, default_parallel, slow, variant, summary, machines: [{machine, fits, text, backend}]}, where status is gpu, gpus, cpu, none or unknown.

Deployments#

Method and path Description
POST $API/deployments Create: {metadata: {name}, spec} (Deployment specification). 201.
GET $API/deployments {items: [...]}, each with its state, replicas, endpoint, plan, gateways and capabilities.
GET $API/deployments/{name} One deployment.
PUT $API/deployments/{name} Change: {spec}, the whole specification. Replicas are replaced one at a time.
DELETE $API/deployments/{name} Delete it and its shares; the model stays. 204.

API keys#

Method and path Description
POST $API/api-keys Create: {metadata: {name}, spec: {deployments?, shared_deployments?, expires_at?}}. 201 with key, shown once.
GET $API/api-keys {items: [...]}: metadata, spec, prefix, created_by, created_at, last_used, expired.
GET $API/api-keys/{name} One key, without its secret.
DELETE $API/api-keys/{name} Revoke it. 204.
PATCH $CONSOLE {"hosted_gateway": true} or false (workspace admins). The workspace then shows hosted_gateway_url.

Shares#

Method and path Description
GET $CONSOLE/share-targets?org=&workspace= Look up a workspace to share with: {org, org_name, workspace, workspace_name, namespace}.
POST $API/deployment-shares Share: {metadata: {name}, spec: {deployment, to: {org, workspace, namespace} \| to_namespace, access: ["call", "playground"], expires_at?, note?}}. 201.
GET $API/deployment-shares {items: [...]}; ?deployment=<name> for one deployment's. Each has status.state Active or Revoked.
GET $API/deployment-shares/{name} One share.
POST $API/deployment-shares/{name}/revoke Revoke it.
DELETE $API/deployment-shares/{name} Delete it. 204.
GET $CONSOLE/shared-deployments Deployments shared with this workspace: {items: [{cluster, share, deployment, owner_namespace, owner, access, expires_at, note, state, kind, model, capabilities}], unreachable}.

The Playground's routes#

The Playground uses these routes; scripts may too (editors and admins).

Method and path Description
POST $API/deployments/{name}/chat Start an answer: an OpenAI chat request (messages, sampling fields, tools, response_format, chat_template_kwargs, reasoning_effort). 202 with {id, state}.
GET $API/deployments/{name}/chat/{id}?after=<offset> What was written since offset: {text, offset, done, state, finish_reason, usage, error, timings}. Read until done.
DELETE $API/deployments/{name}/chat/{id} Stop it.
POST $API/deployments/{name}/embeddings {"input": [...]}, at most 64 texts: OpenAI's embeddings answer.
POST $API/deployments/{name}/transcriptions Speech to text, for a Whisper deployment.
POST $CONSOLE/shared-deployments/{cluster}/{deployment}/chat, …/chat/{id}, …/embeddings, …/transcriptions The same, for a deployment shared with this workspace for the Playground.

model and stream are the deployment's to decide. A request to a deployment at zero wakes it; the answer's state is starting until a replica serves.

Runs that use a model#

A run's specification takes model (see Batch runs that read a model):

Field Type Default Description
spec.model.name string required A model of the run's workspace.
spec.model.mount_path string /model Where its weights are mounted, read-only.
spec.model.engine bool false Run the model's engine beside the run's workers.

Errors#

Code Status Meaning
INVALID_MODEL 400 The model specification is invalid; the message names the field.
MODEL_NOT_FOUND 404 No such model (also when a deployment names one).
MODEL_ALREADY_EXISTS 409 A model of that name exists.
MODEL_IN_USE 409 Live workers use its weights; delete with ?force=true.
MODEL_CONFLICT 409 A drive of yours has the name the model's drive would take.
NO_DATA_LOCATION 409 The machine chosen for a pull has no data location.
NO_ROOM 409 No machine has room for the weights, or none may hold them.
NODE_NOT_FOUND 404 No such machine for this workspace.
CATALOG_ENTRY_NOT_FOUND 404 No library entry by that reference.
CREDENTIAL_REQUIRED 400 A gated library entry needs a credential.
FAMILY_NOT_FOUND 404 No library family by that name.
INVALID_REPO 400 Not owner/name, or an invalid revision.
REPO_GATED 403 The repository is private or gated: add it with its files listed by hand, and a credential.
REPO_NOT_FOUND 404 No such repository or revision on Hugging Face.
HUB_UNREACHABLE 503 Hugging Face did not answer.
INVALID_DEPLOYMENT 400 The deployment specification is invalid, or does not fit its model.
DEPLOYMENT_NOT_FOUND 404 No such deployment.
DEPLOYMENT_ALREADY_EXISTS 409 A deployment of that name exists.
INVALID_API_KEY 400 The key's specification is invalid.
API_KEY_NOT_FOUND 404 No such key.
API_KEY_ALREADY_EXISTS 409 A key of that name exists.
INVALID_DEPLOYMENT_SHARE 400 The share's specification is invalid.
DEPLOYMENT_SHARE_NOT_FOUND, DEPLOYMENT_SHARE_ALREADY_EXISTS 404, 409 —
SHARE_TO_SELF 400 Sharing with the deployment's own workspace.
SHARED_DEPLOYMENT_NOT_FOUND 404 No such deployment is shared with your workspace (or not for this).
NO_READY_BACKEND 503 The deployment has no serving replica and could not be woken in time; retry.
WORKSPACE_ROLE 403 The route is for the workspace's editors.