Skip to content

Data centers#

The rest of the recipes build one thing on a few machines. These build and run a fleet: dozens to hundreds of 8-GPU servers on InfiniBand or RoCE, several teams sharing it, machines that fail and repair themselves, and the record an auditor can check offline. Each recipe is still one real task, start to finish, with the console paths, astra commands and API calls that do it.

  • Bring up a new GPU cluster


    Install the agent on a rack of machines at once, run acceptance, read the Cluster Report, and fix what topology finds before the first real run lands.

    Uses: the installer at scale, join tokens, pools, acceptance, the Cluster Report, topology findings.

  • Adopt Astraeus next to Slurm or Kubernetes


    Install in observe-only mode beside the scheduler you run today, export topology.conf to it, then accept and hand over machines one partition at a time.

    Uses: observe-only mode, the Cluster Report, Slurm's topology.conf, acceptance, astra slurm.

  • Large distributed training


    A gang across many 8-GPU machines, kept on one leaf group's rails, checkpointing to a drive, and resuming on its own after preemption or a lost machine.

    Uses: start: Gang, topology.keep_within, RDMA, checkpoints, on_failure: RestartJob.

  • Self-healing for a production fleet


    An autonomy policy per pool, a BMC power cycle through a neighbour, approvals that page on-call, and a signed receipt for every repair.

    Uses: remediation policies, circuit breakers, BMCs, PagerDuty and Opsgenie alerts, repair receipts, RMA evidence.

  • Share a cluster among teams


    Several pools, several workspaces, fair share and preemption between them, and a reservation that holds a rack for a deadline.

    Uses: pools, quotas, fair-share weights, max priority, reservations.

  • Keep a fleet healthy over time


    Periodic checks on idle machines, a lemon that fails more than its peers, and draining machines for a firmware upgrade without losing work.

    Uses: a pool's periodic checks, lemons, drain and undrain, maintenance.

  • Serve models for the whole company


    One Eos deployment with several replicas, shared with every team's workspace, each with its own API key, behind a gateway at your edge.

    Uses: replicas, scale to zero, sharing a deployment, API keys, the gateway on your machines.

  • Govern and audit


    Single sign-on for everyone, the narrowest role for each person, the audit log for who changed what, and signed repair receipts and RMA evidence an auditor can verify offline.

    Uses: SSO, organisation and workspace roles, guests, the audit log, repair receipts, evidence packs.

Before any of these#

  • You are an owner or admin of the organisation: adding machines, running acceptance, setting pools, quotas, reservations and remediation policies, and reading the audit log are organisation decisions. Roles and permissions has the full list.
  • The API examples use an organisation-scoped token (ASTRALYX_TOKEN), not a workspace one: most of these actions are the organisation's, and its routes refuse a token scoped to one workspace. See API tokens and automation.
  • The rest follows How the recipes are written: specifications are JSON, every step shows Console, then CLI, then API, and <org>, <workspace> and <cluster> stand for yours.