Skip to content

Troubleshooting#

Symptoms an organisation admin meets, grouped by area, each with its cause and fix. Error codes are the code of the API's error answer, also shown by the console. For a run that does not start, see Troubleshooting and FAQ; for a machine, see Machines: troubleshooting.

See the exact reason

The console folds identifiers and exact reasons by default. Turn on Technical details in the account menu (top right) to show them everywhere, for example to copy an error code into a support request.

People and invitations#

Symptom Cause Fix
The invited person never got the e-mail Filtered as spam, or a typo in the address. Check Pending invitations for the address; revoke and invite again. The link cannot be resent.
403 WRONG_ACCOUNT when accepting Signed in with another address than the invited one. Sign out, then sign in or sign up with the invited address.
400 INVALID_TOKEN when accepting The link is older than 7 days, was used, or was revoked. Invite again.
A new member sees no workspace Members see only workspaces they were added to. Add them: workspace Settings → People → Add person.
Add person offers nobody Everyone with the organisation role member is in the workspace already; owners and admins are never offered (they are admins everywhere). Invite the person to the organisation first.
400 NOT_IN_ORG adding someone to a workspace They are not a member of the organisation. Invite them to the organisation first.
The role selector is missing in a row It is your own row, or an owner's and you are not an owner. Ask another admin, or an owner.
409 LAST_OWNER Demoting, removing or leaving as the last owner. Make someone else an owner first.
403 ORG_OWNER_REQUIRED Granting or revoking owner, removing an owner, or deleting the organisation, as an admin. Ask an owner.
Someone removed is a member again They signed in through SSO, which made them a member again. Disable them at the identity provider, then remove them.

Sign-in and SSO#

Symptom Cause Fix
400 DISCOVERY_FAILED saving SSO settings The issuer's discovery document is unreachable, on a private address, or names another issuer. Use the issuer exactly as <issuer>/.well-known/openid-configuration names it.
403 DOMAIN_NOT_ALLOWED The person's e-mail domain is not allowed. Add it to Allowed e-mail domains.
401 SSO_FAILED: the provider has not verified this e-mail No email_verified: true in the ID token. Release that claim to the client at the provider.
401 SSO_FAILED: this account is disabled Astralyx disabled the account. Contact Astralyx support.
404 SSO_NOT_CONFIGURED at /sso Wrong short name, or SSO off. Use the short name in the organisation's addresses (/o/<short name>).
An SSO user cannot sign in with a password Accounts made by SSO have none. Forgot the password? on the sign-in page sets one.
429 TOO_MANY_ATTEMPTS Sign-in attempts are rate-limited. Wait 15 minutes.
403 SIGNUP_CLOSED Sign-up is by invitation. Send an invitation; the person signs up from its link.
The CLI says not signed in or the session is over The CLI session (30 days) expired, or the token was revoked. astra login again, or a new token in ASTRA_TOKEN.

Organisation and workspaces#

Symptom Cause Fix
403 ORG_SUSPENDED Astralyx suspended the organisation; the message says why. Contact Astralyx support.
404 ORG_NOT_FOUND / 404 WORKSPACE_NOT_FOUND for something that exists You are not a member of it (its existence is not revealed). Ask an admin to add you.
403 LIMIT_REACHED creating a workspace The organisation has as many workspaces as its limit. Delete one, or ask Astralyx for more.
409 WORKSPACE_EXISTS The short name is taken in the organisation. Choose another.
403 WORKSPACE_ADMIN_REQUIRED Changing people, the name or retention without being the workspace's admin. Ask a workspace or organisation admin.
A workspace admin cannot change the quota or give access to a cluster Those are organisation decisions. An organisation owner or admin does it.
409 ORG_HAS_CLUSTERS deleting the organisation It still has clusters. Remove the workspaces' access, then the clusters, then delete.

Cluster access and terms#

Symptom Cause Fix
A workspace's work never starts; its Settings shows no cluster It has no access to a cluster. Give access to a cluster.
409 ALREADY_BOUND The workspace already has access to that cluster. Change its terms instead.
409 CHOOSE_CLUSTER adding a first machine The organisation has several clusters. Choose the cluster in the dialog.
503 CLUSTER_REFUSED / 503 CLUSTER_UNREACHABLE The cluster refused or did not answer; nothing changed. Retry; the message gives the cluster's reason.
409 CLUSTER_IN_USE removing a cluster A workspace still has access to it. Remove those accesses first.
403 NAMESPACE_TERMINATING The workspace's access is being removed. Wait, then give access again if needed.
Work waits with Waits for namespace quota: … The workspace holds as much as its quota allows. Wait, or raise the quota.
more than the namespace's quota allows at all One run alone asks for more than the quota. Raise the quota, or ask for less.
403 PRIORITY_NOT_ALLOWED The priority asked is above the workspace's maximum. Lower it, or raise Max priority.
403 HOST_ACCESS_FORBIDDEN A host path or privileged work not granted. Grant it in the terms, or remove it from the spec.
A quota of 0 became unlimited The console shows 0 as empty, and saving makes it unlimited. Set zero limits through the API.
A machine is missing from a workspace's Machines It is outside the workspace's pools. Change the machine's pool, or the workspace's pools.

Machines#

Symptom Cause Fix
403 LIMIT_REACHED on Get the install command Astraeus Cloud's machine limit, counting unused install commands. Remove a machine, wait for unused commands to expire (24 hours), or ask for more.
The dialog keeps Waiting for the machine to connect The install failed, or the machine cannot reach Astralyx over HTTPS. Read the installer's output on the machine; see The machine never appears.
Only admins can add machines Adding machines is an organisation admin action. Ask an owner or admin.
A machine is Down No report for 60 seconds. See The machine is Down.
A machine is Up but takes no work Cordoned, under pressure, reserved, outside the pools, or its GPUs are fenced. Check its page; see Up but takes no work.
Runs needing a drive wait on a new machine It has no data location yet. Choose one on the machine's page.
Update fails The machine could not download or verify the release. See An update fails.
409 NODE_IN_USE removing a machine Work still runs there. Stop the work and remove, or wait.

Usage and cost#

Symptom Cause Fix
Hours shown, cost 0 The cluster has no price for those resources. Set the cluster's Prices.
A GPU model is not priced by gpu:<model> The name differs from what the machine reports. Copy the model name from the hover on GPU-hours, or set a gpu price.
Not counted — the cluster did not answer A cluster could not be read. Its hours are kept; they appear once it answers.
Members see only some workspaces' usage Members see the workspaces they belong to. Owners and admins see every workspace.
403 HOSTED_CLUSTER setting prices Astraeus Cloud's prices are set by Astralyx. —

Alerts and event streams#

Symptom Cause Fix
Machine alerts never arrive The rule is for one workspace. Use every workspace, and the machines.
A delivery is failed The receiver refused, timed out (15 s), or is on a private address. Check Sent; fix the receiver, then Send a test.
Signature mismatch HMAC computed over parsed JSON, or with a hex-decoded key. Use the raw body and the key's ASCII bytes.
A stream is failing The SIEM refused or was unreachable; its status says how. Fix it; the stream resumes where it stopped.
machines events missing from a stream The stream is for one workspace. Make one for every workspace.
400 UNREACHABLE_ADDRESS The URL is not a public address. Use a public endpoint.

Contact support#

Write to [email protected] with the organisation's short name, the time (UTC), what you did, and the error code and message (turn on Technical details to see them). For a machine, add what Machines: troubleshooting asks for.