Govern and audit#
Running a fleet for a company means someone, eventually, asks who could have done this, and how do we know a repair really happened. This recipe connects your identity provider so people sign in through it, sets the narrowest role for each kind of person, gives a contractor exactly one environment and nothing else, reads the audit log for who changed what, and verifies a chain of repair receipts without trusting the console at all.
Before you begin#
- An organisation owner or admin.
- An application you can register at your identity provider (for SSO), or skip that section if you sign in another way.
- An organisation admin's API token,
curlandjq. Below,<org>stands for your organisation's short name, andASTRALYX_APIishttps://api.astralyx.cloud/v1.
1. Connect single sign-on#
Each organisation connects one identity provider over OpenID Connect. Everyone who signs in through it joins your organisation at their first sign-in, with the role you choose.
- Organisation → Single sign-on. Copy the Redirect URI
(
https://console.astralyx.cloud/api/v1/auth/sso/callback). - At your identity provider, register a confidential client with that redirect URI. Note its issuer URL, client ID and client secret.
- Fill in Issuer, Client ID, Client secret, Allowed
e-mail domains (
acme.ai, lab.acme.ai), Role for new members (member), Turn on. - Test in a private window at
https://console.astralyx.cloud/sso.
The provider's ID tokens must carry email_verified: true; without a
verified e-mail, sign-in is refused rather than silently trusted. SSO adds
members; it never removes them — when someone leaves, disable them at the
provider first, then remove them from the organisation so their access
ends at their next request. See
Single sign-on and sign-in.
2. Give each person the narrowest role#
Two layers of roles combine: an organisation role (owner, admin,
member) controls the organisation itself — its people, clusters,
machines and settings; a workspace role (admin, editor, viewer,
auditor) controls what someone may do inside one workspace, in every
product.
| For | Give |
|---|---|
| On-call and managers who only follow work | Workspace viewer |
| The team running runs and models | Workspace editor |
| The team's lead | Workspace admin |
| A compliance reviewer who must not see workloads or data | Workspace auditor — not viewer: a viewer reads logs and data catalogs, an auditor does not |
| The small group who add machines, set quotas and pools, approve repairs | Organisation admin |
| Someone who must never be removed by a lone admin | Organisation owner — keep at least two |
New people are invited with the role they should have
(POST $ASTRALYX_API/organizations/<org>/invitations with
{"to": ["[email protected]"], "workspace": {"slug": "vision", "role": "editor"}});
to change an existing member's role:
$ curl -sS -X PUT "$ASTRALYX_API/organizations/<org>/workspaces/vision/members/$USER_ID" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' -d '{"role": "editor"}'
Organisation owners and admins are already admin in every workspace,
without being added to it: keep that group few, since any of them can grant
host access, which is root on your machines. See
Roles and permissions.
3. Hand a contractor one environment, nothing else#
A vendor engineer needs to look at one development environment to debug a driver issue — not your runs, models, credentials or other environments. Invite them as a guest of that environment, not as a workspace member:
On the environment's People, invite them, choosing Only this environment (guest).
They see that one environment's page, its state and the machine it runs on by name, and may SSH into it and work in it — nothing else of the workspace or the organisation. Take it back from them when the work is done, or make them a member later if they join the team properly. See Roles and permissions: Guests and Invite people.
4. Read the audit log#
The audit log records every platform change — memberships, workspaces, clusters, machines, policies, SSO — with who, what, on which target, from which address.
Organisation → Audit log: When, Who, Action, Target, From; select a row for its details.
$ curl -sS "$ASTRALYX_API/organizations/<org>/audit?limit=5" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
| jq -c '.items[] | {at, user_email, action, target}'
{"at":"2026-10-06T10:41:13Z","user_email":"[email protected]","action":"cluster.node.drain","target":"gpu-03"}
{"at":"2026-10-06T09:12:02Z","user_email":"[email protected]","action":"cluster.remediation.approve","target":"act-20261006-5c2d8a11"}
{"at":"2026-10-05T18:00:11Z","user_email":"[email protected]","action":"cluster.reservation.create","target":"r-3f9a0c1d"}
The audit log is kept 12 months, never less than six — an access
record several jurisdictions' laws require, not only a convenience. It
covers organisation, workspace and cluster changes
(cluster.node.*, cluster.remediation.*, cluster.reservation.*,
cluster.validation.*, sso.*, workspace.*…); day-to-day work inside a
workspace (a run submitted, a drive deleted) is not a platform change —
it appears as an event and in the
cluster's own request log, kept by Astraeus but not shown in the console.
See Events, audit and event streams.
5. Verify repair receipts, offline#
Every repair self-healing carried out — automatic, policy-driven, or approved by a person — left a signed, chained receipt. An auditor with no access to Astraeus at all can still check that the chain is intact and every signature is genuine, against the cluster's own public keys:
$ astra astraeus repair-receipts --limit 50 --json > receipts.json
$ astra astraeus repair-receipts --jwks cluster-jwks.json --verify
10 ok
11 ok
12 ok
Without network access to fetch the keys live, save them once
(GET /cluster-identity/jwks) and pass --jwks. Each receipt names what
was decided, by whom (a person, or policy), why, and how it ended — the
same record Self-healing for a production fleet
produces as it runs. A machine replaced along the way carries its own
RMA evidence pack, also signed and independently verifiable:
See Receipts and RMA evidence.
6. Optional: send every event to your SIEM#
For a continuous feed rather than point-in-time checks, an event stream sends every event of the categories you choose — machines, outbound denials, self-healing, and more — to a webhook, Splunk's HTTP Event Collector, or syslog over TLS, mapped to OCSF 1.3.0:
$ curl -sS -X POST "$ASTRALYX_API/organizations/<org>/event-streams" \
-H "Authorization: Bearer $ASTRALYX_TOKEN" -H "Content-Type: application/json" \
-d '{"kind": "splunk_hec", "name": "splunk", "url": "https://<splunk-host>:8088/services/collector/event",
"token": "<HEC token>", "categories": ["machines", "outbound_denials", "approvals"], "workspace": null}'
See Event streams.
What you get#
- Everyone signing in through your own identity provider, with access that ends the moment you remove them from the organisation.
- A contractor who reaches exactly one environment and nothing else of your fleet.
- An audit log that answers "who changed this" for a year, and a chain of signed repair receipts an auditor can verify with no access to Astraeus at all.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
400 DISCOVERY_FAILED saving SSO |
The discovery document is unreachable, or names another issuer. | Use the issuer exactly as <issuer>/.well-known/openid-configuration names it. |
403 DOMAIN_NOT_ALLOWED at sign-in |
The e-mail's domain is not in Allowed e-mail domains. | Add the domain, or empty the list. |
| Someone removed from the organisation signs in again | They were still enabled at the identity provider; SSO re-added them. | Disable them at the provider first, then remove them here. |
| A guest sees more than one environment | They were added as a workspace member, not a guest. | Remove their membership; invite them to the one environment as a guest instead. |
403 GUEST_FORBIDDEN |
A guest asked for anything beyond what was given to them. | Expected: make them a member if they need more. |
astra astraeus repair-receipts --verify fails on one entry |
The chain was broken (an edited export), or the wrong cluster's keys were used. | Re-fetch --jwks from the right cluster; re-export the receipts. |