Monitor a deployment#
A deployment's Metrics tab shows how it serves: the same measures whichever engine runs it (vLLM or llama.cpp), each labelled with where it comes from — the engine, the replicas' machines, or the gateways. This page lists what is measured, how to read it live and over time, and how to be alerted.
What is measured#
| Measure | vLLM | llama.cpp |
|---|---|---|
| Tokens generated per second | The engine | The engine |
| Prompt tokens processed per second | The engine | The engine |
| Requests per second | Requests finished (the engine) | The gateways' count: llama.cpp counts no requests |
| Time to first token, p50 and p95 | The engine | Not available: llama.cpp keeps no histograms |
| Time between tokens | p50 (the engine) | Mean: seconds spent generating over tokens generated |
| Request latency, p50 and p95 | The engine | How long clients waited at the gateways |
| Generation speed | — | llama.cpp's own mean, tokens per second while generating |
| Requests running and waiting | Running, waiting | Slots busy, requests deferred |
| Preemptions per second | The engine | — |
| KV cache in use | The engine, 0–100 % | The engine, 0–100 % |
| GPU utilisation, GPU memory used and total, power, temperature | Each replica's GPUs | Each replica's GPUs |
| CPU and memory | Each replica | Each replica |
Gateways (the gateway on your machines and the hosted one):
- Requests per second by status:
2xx,4xx(the request was refused, or a limit was reached) and5xx(the model failed, could not be reached, or did not answer in time). - Latency as clients saw it, p50 and p95: from the request arriving to the last byte of the answer (for a streamed answer, the whole stream).
- Tokens per second by API key. Calls from another workspace under a
share show as
shared:<their workspace's namespace>, not under their key's name.
Only counts and durations are recorded: never a prompt or an answer.
Cold starts: how many times a request woke the deployment at zero in the last 24 hours (7 days with that range), and how long it took until a replica served — mean and longest.
Read it in the console#
- Open Eos → Deployments and pick the deployment.
- Open the Metrics tab. It opens Live:
- a row of the numbers that matter now: tokens per second, requests
running and waiting, time to the first token (request latency for
llama.cpp), KV cache, GPU utilisation and memory, gateway requests
per second, and errors per second (
4xx·5xx); - charts that move every 2 seconds over the last 15 minutes, with updated N s ago beside them.
- a row of the numbers that matter now: tokens per second, requests
running and waiting, time to the first token (request latency for
llama.cpp), KV cache, GPU utilisation and memory, gateway requests
per second, and errors per second (
- Choose 15 min, 1 hour, 24 hours or 7 days to see the history instead (refreshed every 30 seconds); Live returns to live updates.
- All replicas adds the replicas up (rates and memory are summed; utilisation and KV cache averaged; quantiles are over every replica's requests). Per replica draws one line per replica.
A hidden browser tab stops live updates until it is shown again. In live mode, GPU power, temperature and total memory are in the history only: choose a range to see them.
An empty chart says why:
| Message | Meaning |
|---|---|
| no replica runs (scaled to zero) | The deployment is at zero: nothing runs, so neither the engine nor its machine reports. The next request starts a replica. |
| the engine has not reported this yet | The replica is still loading the model, or the engine does not expose that measure. |
| no requests through the gateways in this range | No program called it through a gateway. The Playground is not counted here. |
Watch it from the command line#
astra eos top chat # redrawn every 2 seconds; Ctrl-C to quit
astra eos top chat --once # print one reading and exit
The first lines are the deployment's throughput, latency, queue, KV cache and gateway numbers; then one line per replica: its machine, whether it serves, tokens per second, requests running and waiting, KV cache, GPU, GPU memory, and how long ago its machine last reported. A measure the engine does not have is left out.
Read it from the API#
curl -H "Authorization: Bearer $ASTRALYX_TOKEN" \
"https://api.astralyx.cloud/v1/eos/deployments/chat/metrics?range=1h&per=replica"
| Parameter | Values | Default |
|---|---|---|
range |
15m, 1h, 6h, 24h, 7d |
1h |
per |
total (the deployment's), replica (one series per replica) |
total |
measures |
Measure ids, comma-separated (generation_tokens_per_second,ttft_p95) |
all |
Each measure comes with its unit, its source, the series it was read from
and its points ([Unix milliseconds, value], about 240 per range).
GET …/eos/deployments/chat/metrics/live streams server-sent events: one
measures event, then a reading every 2 seconds with each measure's
value now, per replica, and each replica's last report. See the
Eos API reference.
Get alerted#
Alert rules (Organisation → Alerts) fire on events; see Alerts. Two rules that watch deployments:
- When a deployment goes down: a rule for
deployment_downin the deployment's workspace, sent to your team's Slack channel. It fires when a deployment failed, or had no replica serving for 5 minutes after serving, while not scaled to zero on purpose. - When its machines are in trouble: a rule for
machine_downandgpu_degradedin every workspace, sent to a webhook. It fires when a machine stops answering, or its GPUs are not healthy (a fault, no driver).
Alert rules do not watch thresholds on these numbers. To be told when the
error rate or the latency is too high, read GET …/metrics?range=15m from
your own monitoring every minute and compare, for example, the 5xx series
of gateway_requests to its total, or the last point of
gateway_latency_p95.