Skip to content

Goodput#

GPUs are the expensive part of a cluster, and a run holding them is not the same as a run using them. Astraeus counts, every minute, where the GPU time each run asked for went, and adds it up per run, machine, pool and workspace, by the hour and by the week. Goodput is the share of the GPU time held that was used productively. This page explains the categories, how to read them, and how to get them into your own monitoring.

The categories#

Every minute, each worker of every run still going is put in one category, for the time since the last count, times the GPUs it holds (or asked for, while it waits):

Category The worker was Held?
productive running, its GPUs busy (5% or more); or winding down when its owner asked (a stop, its time limit) yes
idle running, its GPUs idle yes
stalled running while stalled (from when it stopped making progress); stopping because it stalled; waiting to be placed again after held while placed
starting placed, not yet running, on its first attempt: its image, its drives, its gang's barrier yes
restart the same on a later attempt: the cost of restarting yes
failure failed or lost and still holding its GPUs; waiting to be placed again after a failure, a lost machine or a GPU fault held while placed
preemption stopping for higher-priority work or a drain; waiting to be placed again after held while placed
waiting waiting for its first placement no
  • Allocated (held) is every category but the waits for placement.
  • Goodput = productive ÷ allocated.
  • Lost is everything asked for but productive, waits included.

Astraeus does not know what your run computes. Work an attempt did and your code did not keep (on a drive) when it failed is not counted as lost here: it was productive time then. What is counted is time — what failures and preemptions cost before the run was working again, GPUs held and not used, and stalls.

Only GPU time is counted. Runs without GPUs have no goodput.

See it#

  • Goodput in the workspace's Astraeus menu: goodput, GPU time held, productive, not productive (with the biggest loss) and waiting, then a row per run, machine, pool or workspace with a bar of where its time went. By week (8) shows the last eight weeks. Workers stalled and stragglers now, and threshold alerts firing, are listed first.
  • A run's page shows its own goodput and bar under Goodput.
$ astra astraeus goodput --by run --days 7
GPU-hours
RUN                              HELD     USED GOODPUT    IDLE STALLED   START RESTART FAILURE PREEMPT WAITING
llm                             323.1    300.1     93%     0.0     6.2     2.1     4.1    12.7     0.0     0.0
sweep-7                         146.5    112.4     77%    21.0     0.0     7.7     0.0     0.0     3.3    18.0
TOTAL                           469.6    412.5     88%    21.0     6.2     9.8     4.1    12.7     3.3    18.0

USED = productive (running, GPUs busy); GOODPUT = used / held; WAITING = asked for, not yet held
$ astra astraeus goodput --weeks 8 --by pool
$ astra astraeus goodput --run llm
$ curl -sS "$ASTRALYX_API/goodput?by=machine&start=2026-10-01T00:00:00Z" -H "Authorization: Bearer $ASTRALYX_TOKEN"
$ curl -sS "$ASTRALYX_API/goodput/weekly?weeks=8&by=run" -H "Authorization: Bearer $ASTRALYX_TOKEN"
$ curl -sS "$ASTRALYX_API/runs/llm/goodput" -H "Authorization: Bearer $ASTRALYX_TOKEN"

Each ledger is GPU-hours by category (gpu_hours), allocated_gpu_hours, productive_gpu_hours, lost_gpu_hours, waiting_gpu_hours and goodput (absent when nothing was held). by is workspace (default), run, machine or pool; a period is at most 400 days.

Within a workspace you see your own runs' time (on the machines they used). Your organisation's admins see every workspace's and machine's of the organisation on a cluster (/organizations/<org>/clusters/<cluster>/goodput).

Prometheus and OpenTelemetry#

GET /goodput/metrics serves the same as Prometheus' text exposition, for a Prometheus server or an OpenTelemetry Collector's prometheus receiver to scrape with your API token:

otel-collector.yaml
receivers:
  prometheus:
    config:
      scrape_configs:
        - job_name: astralyx-goodput
          scrape_interval: 60s
          scheme: https
          metrics_path: /v1/goodput/metrics
          authorization:
            credentials: ${env:ASTRALYX_TOKEN}
          static_configs:
            - targets: [api.astralyx.cloud]
Metric Labels Description
astraeus_goodput_ratio scope (cluster, workspace, machine, pool), name, window (hour, day) Productive over held, in the current UTC hour or the last 24 hours.
astraeus_goodput_gpu_hours scope, name, window, category GPU-hours by category.
astraeus_run_goodput_gpu_hours namespace, run, category A run's, since it was submitted (runs still going).
astraeus_run_goodput_ratio namespace, run Its goodput.
astraeus_workers_stalled namespace Workers stalled now.
astraeus_runs_with_straggler namespace Runs whose gang has a straggler now.
astraeus_threshold_alerts_firing metric Threshold alerts firing (a cluster's operators only).

All are gauges over whole hours (the hourly counts are kept 400 days and pruned, so a counter would go down).

Alerts on goodput#

A workspace's or pool's goodput over the last complete hour under 50%, and 20 points or more under the day before, fires the goodput_drop alert (at least one GPU-hour held in that hour). The thresholds are the cluster's; see Threshold alerts.