Skip to content

Errors#

The Astraeus API at https://console.astralyx.cloud/api/v1 — including a workspace's cluster path — answers a refused or failed request with an HTTP status and a JSON body that names the error with a stable code. This page describes that body and lists the codes, grouped by area, with their meaning and what to do.

The error response#

{
  "code": "PRIORITY_NOT_ALLOWED",
  "message": "priority 50 is above namespace ws-3f9a1c07b2e4's maximum of 0",
  "detail": {"name": "train"}
}
Field Type Description
code string Stable, machine-readable, UPPER_SNAKE_CASE. Match on this.
message string Human-readable. It may change between releases; do not parse it.
detail object Optional context: the object's name, the offending field, and so on. Omitted when empty.

A validation failure lists every bad field in detail, by field path:

{
  "code": "VALIDATION_ERROR",
  "message": "Failed to validate object",
  "detail": {
    "spec.priority": "Must be between -1000 and 1000, got 5000",
    "spec.array.size": "Must be between 1 and 10000, got 0"
  }
}

Under a workspace's cluster path, names in errors are shown as the workspace writes them (train), without the namespace. For the console's own operations on a cluster (giving a workspace access, cordoning a machine), a 400, 403, 404 or 409 from the cluster keeps its code and message; any other refusal becomes 503 with "the cluster answered …".

Statuses#

Status Meaning Retry?
400 The request is malformed or invalid. No: fix the request.
401 Missing, unknown or expired credentials. After signing in again.
403 Authenticated, but not allowed. No, unless a role or setting changes.
404 Not found — or not visible to you: organisations, workspaces and machines outside your reach answer 404, not 403. No.
409 Already exists, or a state conflict (in use, changed concurrently). Conflicts from concurrent changes: re-read and retry.
429 Too many requests. After the window, or after Retry-After.
500 A server fault. Yes, with backoff.
502 An upstream dependency failed. Yes, with backoff.
503 Temporarily unavailable: a cluster or provider did not answer, or the server is at capacity. Yes, with backoff; honour Retry-After.
504 An upstream dependency timed out (for example a machine asked for a log). Yes.

Server faults are redacted: a 500 always reads {"code": "INTERNAL_SERVER_ERROR", "message": "Internal server error"}, a 502 BAD_GATEWAY, a 504 TIMEOUT ("A timeout occurred"), with no detail. Astralyx sees the full error.

Errors outside this format#

  • Some requests are refused by the HTTP layer before Astraeus reads them, with the standard status and a plain-text body: a JSON body over 2 MiB on most endpoints (413), a body that is not valid JSON or lacks a required field (400 or 422), a missing Content-Type: application/json (415), a method a route does not have (405).
  • An unknown path under /api/ answers 404 {"code": "NOT_FOUND", "message": "no such route"}; under a workspace's cluster path, an unknown path answers 404 with an empty body.
  • The CLI sign-in endpoint POST /api/v1/auth/device/token answers in the OAuth device-flow form: 400 {"error": "authorization_pending"} (keep polling), "expired_token" (start again) or "invalid_grant" (unknown or already used).
  • A watch stream that falls behind receives an ERROR event and should list again; see REST API.

Request format#

Code Status Meaning What to do
VALIDATION_ERROR 400 One or more fields are invalid; detail lists them. Fix each field.
UNKNOWN_FIELD 400 A run's spec has fields the API does not know (often a typo). Remove or correct them; see Run specification.
INVALID_JOB 400 The run's body could not be read as a run. Check the JSON against the run specification.
INVALID_JSON 400 The body is not valid JSON, or not a JSON object. Send a JSON object.
BODY_TOO_LARGE 400 The body is over 16 MiB (under a workspace's cluster path). Send less; put data on a drive.
INVALID_NAME 400 A name contains ., is empty, or is otherwise not a valid name. Inside a workspace, names never contain .. Use lower-case letters, digits and -.
NAME_NOT_DNS_SAFE 400 A run, replica group or endpoint name cannot be published in cluster DNS; detail.dns_name shows the derived name. Shorten the name or use only letters, digits, - and _.
INVALID_ORDER 400 order is not asc or desc. Fix the parameter.
INVALID_SORT_BY 400 Unknown sort_by for workers. Use created_at, started_at, finished_at, name or updated_at.
INVALID_STATE, INVALID_TASK_STATE 400 A state filter or value is not a valid state. See Run and worker states.
INVALID_METHOD 400 An HTTP method that cannot be forwarded to a worker or endpoint. Use a standard method.
INVALID_RESOURCE_VERSION 400 A schedule's resourceVersion is not a number. Re-read the schedule.
CONFLICT 409 A worker lost a concurrent write. Retry.

Authentication and authorization#

Console#

Code Status Meaning What to do
UNAUTHENTICATED 401 No session cookie or bearer token, or it is unknown, expired, or its account disabled. Sign in again, or use a valid API token.
CSRF 403 A request authenticated by the session cookie changed something without the X-Astraeus-Client header. Send X-Astraeus-Client (any value), or use a bearer token.
INVALID_CREDENTIALS 401 / 403 Sign-in: wrong e-mail or password, or the account is not verified yet (401). Password change: the current password is wrong (403). Check the credentials; verify the e-mail; reset the password.
TOO_MANY_ATTEMPTS 429 Too many sign-ins, sign-ups, reset or device requests. Wait; see Limits.
SIGNUP_CLOSED 403 Sign-up is by invitation. Ask for an invitation.
INVALID_EMAIL 400 Not a valid e-mail address. Fix it.
WEAK_PASSWORD 400 Shorter than 12 characters, or over 1024 bytes. Choose another.
INVALID_TOKEN 400 A verification, reset or invitation link is invalid, used or expired. Ask for a new one.
WRONG_ACCOUNT 403 The invitation is for another e-mail address. Sign in with the invited address.
CODE_NOT_FOUND 404 No pending CLI sign-in with that code (expired after 10 minutes, or already approved). Run astra login again.
TOKEN_NOT_FOUND 404 No such API token of yours. List tokens with GET /me/tokens.
SSO_FAILED 401 Single sign-on failed; the message says why (provider refused, ID token invalid, unverified e-mail, sign-in too slow, account disabled). See SSO troubleshooting.
SSO_NOT_CONFIGURED 404 The organisation has no single sign-on. Check the organisation's short name.
DOMAIN_NOT_ALLOWED 403 The e-mail's domain is not allowed by the organisation's SSO. Ask an admin to allow it.
DISCOVERY_FAILED 400 Saving SSO: the issuer's discovery document could not be read, or names another issuer. Check the issuer URL.
INVALID_ISSUER 400 The issuer is not an http(s) URL. Fix it.
CLIENT_SECRET_REQUIRED 400 First SSO configuration without a client secret. Send client_secret.
PROVIDER_UNREACHABLE, PROVIDER_INVALID 503 The identity provider did not answer, or published an unusable address. Retry; check the provider.
ORG_SUSPENDED 403 Astralyx suspended the organisation. Write to [email protected].

Inside a cluster#

Code Status Meaning What to do
UNAUTHORIZED 401 Missing bearer token, unknown token, or an unused token past its enrollment window. Use a valid token.
NODE_REMOVED 401 The machine was removed from the cluster; its old credentials are refused. Install the machine again.
AUTH_BACKEND_UNAVAILABLE 503 The credentials could not be checked right now. Retry.
RBAC_FORBIDDEN 403 The role may not do this: role viewer may not create jobs in ws-…. Ask for a role that allows it; see Roles and permissions.
NAMESPACE_FORBIDDEN 403 The path or body reaches outside the workspace: another namespace's object, a reference into it, or something not available inside a workspace. Name only the workspace's own objects.
HOST_ACCESS_FORBIDDEN 403 The spec uses a host path or a privilege the workspace was not granted; the message names each field. See host access.
PRIORITY_NOT_ALLOWED 403 The run's priority is above the workspace's maximum. Lower it, or ask for a higher max_priority.
NAMESPACE_TERMINATING 403 The workspace's access to this cluster is being removed; no writes. Wait.
NAMESPACE_NOT_FOUND 404 The workspace has no namespace on this cluster. Give the workspace access to the cluster.
TENANT_FORBIDDEN, TENANT_MISMATCH 403 A request for an organisation reached something not that organisation's. Report it to [email protected]: it indicates an error on Astralyx's side.
NODE_NAME_TAKEN 403 A machine tried to register under a name another machine holds. Install with another --name.
RATE_LIMITED 429 A machine exceeded its request rate (burst 60, 5 per second). Retry-After: 5. The agent backs off by itself.
OVERLOADED 503 The cluster is at capacity for the moment. Retry-After: 1. Retry after a second.

Organisations, workspaces and clusters#

Code Status Meaning What to do
ORG_NOT_FOUND 404 No such organisation, or you are not a member. Check the slug and your membership.
ORG_EXISTS 409 The slug is taken. Choose another.
ORG_ADMIN_REQUIRED 403 Owners and admins only. Ask one.
ORG_OWNER_REQUIRED 403 Owners only (ownership changes, deleting the organisation). Ask an owner.
ORG_HAS_CLUSTERS 409 An organisation with clusters cannot be deleted. Write to [email protected].
INVALID_SLUG 400 Slugs are 2–40 lower-case letters, digits and -, not starting or ending with -. Fix it.
MEMBER_NOT_FOUND 404 The user is not a member. —
LAST_OWNER 409 The organisation would have no owner. Make another owner first.
INVALID_ROLE 400 Unknown role, or a role that cannot be used there (invitations: admin, member; SSO: member, admin). Use a valid role.
INVITATION_NOT_FOUND 404 No pending invitation with that ID. —
NOT_IN_ORG 400 Adding to a workspace someone outside the organisation. Invite them to the organisation first.
WORKSPACE_NOT_FOUND 404 No such workspace, or you are not in it. Check the slug and your membership.
WORKSPACE_EXISTS 409 The workspace slug is taken in the organisation. Choose another.
WORKSPACE_ADMIN_REQUIRED 403 Workspace admins only. Ask one.
WORKSPACE_EDITOR_REQUIRED 403 Saving or deleting run templates needs editor. —
INVALID_TEMPLATE 400 A run template needs a name (≤ 100 characters) and a spec object (≤ 256 KiB). Fix it.
TEMPLATE_NOT_FOUND 404 No such run template. —
LIMIT_REACHED 403 The organisation has as many workspaces, or Astraeus Cloud machines, as allowed. Remove one, or ask Astralyx at [email protected].
ALREADY_BOUND 409 The workspace already has access to that cluster. Change the terms instead.
NOT_BOUND 404 The workspace has no access to that cluster. Give it access.
CHOOSE_CLUSTER 409 Connecting a workspace: the organisation has several clusters. Name the cluster.
CLUSTER_NOT_FOUND 404 No such cluster in the organisation. —
CLUSTER_REFUSED 503 Giving a workspace access: the cluster refused it; the message says why. Read the message; retry.
CLUSTER_UNREACHABLE 503 The cluster did not answer. Retry; if it persists, write to [email protected].
UNREACHABLE_ADDRESS 400 An address you gave (an identity provider, a webhook, an event stream) is not public, so Astralyx will not reach it. Use a public address.
HOSTED_CLUSTER 400 / 403 Astraeus Cloud's settings and prices are set by Astralyx. —
RESPONSE_TOO_LARGE 503 / 400 A response over 256 MiB under a workspace's cluster path (503), or a machine's answer over 1 MiB (400). Ask for less (filters, tail).
INVALID_PERIOD 400 Usage: start not before end, or more than 400 days. Narrow the period.
INVALID_PRICE 400 A price table entry is unknown or negative. See prices.
INVALID_RESERVATION 400 A reservation names neither machines nor a pool. Name one.
INVALID_LABEL 400 A pool, site, rack or fabric value: letters, digits and -._:/, at most 63. Fix it.
INVALID_PATH 400 A data location must be an absolute path. Fix it.

Alerts, events and audit#

Code Status Meaning
INVALID_CHANNEL 400 Bad alert channel: kind, addresses or URL.
CHANNEL_NOT_FOUND 404 No such alert channel.
INVALID_RULE 400 An alert rule without events, or with an unknown one.
RULE_NOT_FOUND 404 No such alert rule.
INVALID_EVENT_STREAM 400 Bad event stream: kind, address, token, CA or categories.
EVENT_STREAM_NOT_FOUND 404 No such event stream.
INVALID_RETENTION 400 Event retention is 1 to 3660 days, or null.

Runs and workers#

Code Status Meaning What to do
JOB_NOT_FOUND 404 No such run. —
JOB_ALREADY_EXISTS 409 A run with that name exists. Choose another name, or delete the old run.
MODELS_NOT_SERVED 400 The run names a model and this cluster does not resolve models yet. Mount the weights from a drive.
TASK_NOT_FOUND 404 No such worker. —
TASK_ALREADY_EXISTS 409 A worker with that name exists. —
TASK_HAS_NO_JOB 400 A worker without its run. —
TASK_NOT_PLACED 409 The worker is not on a machine yet: no log, trace or activity. Wait until it is placed.
TASK_LOG_UNAVAILABLE, TASK_TRACE_UNAVAILABLE, TASK_ACTIVITY_UNAVAILABLE 400 The machine could not return them; the message says why. —
TASK_HAS_NO_TOOLS 400 The worker calls no tools: it has no trace. —
INVALID_SINCE 400 since is not an RFC 3339 time. Fix it.
TIMEOUT 504 The machine did not answer within 30 s (logs, trace, activity). Retry; check the machine.
CRONJOB_NOT_FOUND 404 No such schedule. —
CRONJOB_ALREADY_EXISTS 409 A schedule with that name exists. —
INVALID_SCHEDULE 400 The schedule's timing is invalid, or has no next run; the message says why. Fix it; see Schedules.
CRONJOB_NOT_ACTIVE 400 Triggering a paused or completed schedule. Resume it first.
CRONJOB_MAX_RUNS_REACHED 400 The schedule reached its max_runs. —
CRONJOB_TRIGGER_CONFLICT 409 Another trigger is in flight. Retry.

Machines and reservations#

Code Status Meaning What to do
NODE_NOT_FOUND 404 No such machine — or, in a workspace, not in its pools. Check the name and the pools.
NODE_ALREADY_EXISTS 409 A machine with that name is registered. —
NODE_CONFLICT 409 The machine changed concurrently. Retry.
NODE_IN_USE 409 Removing a machine that still runs live workers. Drain it, or remove with force.
NODE_HOLDS_DRIVE_COPIES 409 The machine holds drive copies under its data location. Delete those drives first.
INVALID_NODE 400 Invalid machine record. —
RESERVED_LABEL 400 Labels under astraeus.io are set by the system. Use another key.
ENROLLMENT_NOT_FOUND 404 No unused enrollment token with that ID. —
ENROLLMENT_USED 401 The enrollment token was used by another machine. Make a new one.
INVALID_TTL 400 Enrollment token validity is 60 to 604800 s. Fix it.
CLEAR_FAULT_FAILED 400 The machine could not clear the GPU fault. Read the message.
INVALID_VERSION 400 An upgrade's release name has characters other than letters, digits and ._-. Fix it.
LIVE_UNAVAILABLE 503 Live metrics of the machine are not available. Retry.
METRICS_NOT_FOUND 404 No metrics received for that machine, run or worker. Wait for its first report.
INVALID_QUERY 400 A series query is invalid. See Logs, metrics and debugging.
TSDB_NOT_CONFIGURED, TSDB_UNAVAILABLE 503 Metric series are not available right now. Retry.
NODE_RESERVATION_NOT_FOUND 404 No such reservation. —
NODE_RESERVATION_ALREADY_EXISTS 409 A reservation with that name exists. —
INVALID_NODE_RESERVATION 400 Invalid reservation (window, machines, pool). Read the message.
WORKLOAD_IDENTITY_NOT_CONFIGURED 404 The cluster issues no worker identities. —

Data, credentials and services#

Code Status Meaning What to do
DATAVOLUME_NOT_FOUND 404 No such drive. —
DATAVOLUME_ALREADY_EXISTS 409 A drive with that name exists. —
DATAVOLUME_IN_USE 409 Live workers use the drive. Stop them first.
DATAVOLUME_MANAGED 409 The drive is managed by another object. Delete that object instead.
INVALID_DATAVOLUME 400 Invalid drive spec. Read the message.
CONNECTOR_NOT_FOUND 404 No such data source. —
CONNECTOR_ALREADY_EXISTS 409 A data source with that name exists. —
INVALID_CONNECTOR 400 Invalid data source spec. Read the message.
CREDENTIALS_NOT_ACCEPTED 400 The data source carries a secret value. Reference a credential instead.
INDEX_NOT_READY 404 The data source has not been indexed yet. Wait.
EXTERNAL_SECRET_NOT_FOUND 404 No such credential. —
EXTERNAL_SECRET_ALREADY_EXISTS 409 A credential with that name exists. —
INVALID_EXTERNAL_SECRET 400 Invalid credential spec. Read the message.
ENDPOINT_NOT_FOUND 404 No such endpoint. —
ENDPOINT_ALREADY_EXISTS 409 An endpoint with that name exists. —
INVALID_ENDPOINT 400 Invalid endpoint spec. Read the message.
MANAGED_BY_JOB 409 The endpoint is a run's external access. Change the run instead.
PORT_TAKEN 409 The listen port is reserved by another endpoint. Choose another port.
PORTS_EXHAUSTED 409 No listen port left in 30000–32767. Delete unused endpoints.
SCALING_GROUP_NOT_FOUND 404 No such replica group. —
SCALING_GROUP_ALREADY_EXISTS 409 A replica group with that name exists. —
SCALING_GROUP_MANAGED 409 The replica group is managed by another object. Change that object instead.
INVALID_SCALING_GROUP 400 Invalid replica group spec. Read the message.

Errors of Eos (models, deployments, API keys), Anemoi (agents, flows, approvals, budgets, evaluations, evidence) and Hesperus (notebooks) follow the same format; see their documentation.