Goodput#
GPUs are the expensive part of a cluster, and a run holding them is not the same as a run using them. Astraeus counts, every minute, where the GPU time each run asked for went, and adds it up per run, machine, pool and workspace, by the hour and by the week. Goodput is the share of the GPU time held that was used productively. This page explains the categories, how to read them, and how to get them into your own monitoring.
The categories#
Every minute, each worker of every run still going is put in one category, for the time since the last count, times the GPUs it holds (or asked for, while it waits):
| Category | The worker was | Held? |
|---|---|---|
| productive | running, its GPUs busy (5% or more); or winding down when its owner asked (a stop, its time limit) | yes |
| idle | running, its GPUs idle | yes |
| stalled | running while stalled (from when it stopped making progress); stopping because it stalled; waiting to be placed again after | held while placed |
| starting | placed, not yet running, on its first attempt: its image, its drives, its gang's barrier | yes |
| restart | the same on a later attempt: the cost of restarting | yes |
| failure | failed or lost and still holding its GPUs; waiting to be placed again after a failure, a lost machine or a GPU fault | held while placed |
| preemption | stopping for higher-priority work or a drain; waiting to be placed again after | held while placed |
| waiting | waiting for its first placement | no |
- Allocated (held) is every category but the waits for placement.
- Goodput = productive ÷ allocated.
- Lost is everything asked for but productive, waits included.
Astraeus does not know what your run computes. Work an attempt did and your code did not keep (on a drive) when it failed is not counted as lost here: it was productive time then. What is counted is time — what failures and preemptions cost before the run was working again, GPUs held and not used, and stalls.
Only GPU time is counted. Runs without GPUs have no goodput.
See it#
- Goodput in the workspace's Astraeus menu: goodput, GPU time held, productive, not productive (with the biggest loss) and waiting, then a row per run, machine, pool or workspace with a bar of where its time went. By week (8) shows the last eight weeks. Workers stalled and stragglers now, and threshold alerts firing, are listed first.
- A run's page shows its own goodput and bar under Goodput.
$ astra astraeus goodput --by run --days 7
GPU-hours
RUN HELD USED GOODPUT IDLE STALLED START RESTART FAILURE PREEMPT WAITING
llm 323.1 300.1 93% 0.0 6.2 2.1 4.1 12.7 0.0 0.0
sweep-7 146.5 112.4 77% 21.0 0.0 7.7 0.0 0.0 3.3 18.0
TOTAL 469.6 412.5 88% 21.0 6.2 9.8 4.1 12.7 3.3 18.0
USED = productive (running, GPUs busy); GOODPUT = used / held; WAITING = asked for, not yet held
$ astra astraeus goodput --weeks 8 --by pool
$ astra astraeus goodput --run llm
$ curl -sS "$ASTRALYX_API/goodput?by=machine&start=2026-10-01T00:00:00Z" -H "Authorization: Bearer $ASTRALYX_TOKEN"
$ curl -sS "$ASTRALYX_API/goodput/weekly?weeks=8&by=run" -H "Authorization: Bearer $ASTRALYX_TOKEN"
$ curl -sS "$ASTRALYX_API/runs/llm/goodput" -H "Authorization: Bearer $ASTRALYX_TOKEN"
Each ledger is GPU-hours by category (gpu_hours), allocated_gpu_hours, productive_gpu_hours, lost_gpu_hours, waiting_gpu_hours and goodput (absent when nothing was held). by is workspace (default), run, machine or pool; a period is at most 400 days.
Within a workspace you see your own runs' time (on the machines they used). Your organisation's admins see every workspace's and machine's of the organisation on a cluster (/organizations/<org>/clusters/<cluster>/goodput).
Prometheus and OpenTelemetry#
GET /goodput/metrics serves the same as Prometheus' text exposition, for a Prometheus server or an OpenTelemetry Collector's prometheus receiver to scrape with your API token:
receivers:
prometheus:
config:
scrape_configs:
- job_name: astralyx-goodput
scrape_interval: 60s
scheme: https
metrics_path: /v1/goodput/metrics
authorization:
credentials: ${env:ASTRALYX_TOKEN}
static_configs:
- targets: [api.astralyx.cloud]
| Metric | Labels | Description |
|---|---|---|
astraeus_goodput_ratio |
scope (cluster, workspace, machine, pool), name, window (hour, day) |
Productive over held, in the current UTC hour or the last 24 hours. |
astraeus_goodput_gpu_hours |
scope, name, window, category |
GPU-hours by category. |
astraeus_run_goodput_gpu_hours |
namespace, run, category |
A run's, since it was submitted (runs still going). |
astraeus_run_goodput_ratio |
namespace, run |
Its goodput. |
astraeus_workers_stalled |
namespace |
Workers stalled now. |
astraeus_runs_with_straggler |
namespace |
Runs whose gang has a straggler now. |
astraeus_threshold_alerts_firing |
metric |
Threshold alerts firing (a cluster's operators only). |
All are gauges over whole hours (the hourly counts are kept 400 days and pruned, so a counter would go down).
Alerts on goodput#
A workspace's or pool's goodput over the last complete hour under 50%, and 20 points or more under the day before, fires the goodput_drop alert (at least one GPU-hour held in that hour). The thresholds are the cluster's; see Threshold alerts.