Skip to content

Monitor a deployment#

A deployment's Metrics tab shows how it serves: the same measures whichever engine runs it (vLLM or llama.cpp), each labelled with where it comes from — the engine, the replicas' machines, or the gateways. This page lists what is measured, how to read it live and over time, and how to be alerted.

What is measured#

Measure vLLM llama.cpp
Tokens generated per second The engine The engine
Prompt tokens processed per second The engine The engine
Requests per second Requests finished (the engine) The gateways' count: llama.cpp counts no requests
Time to first token, p50 and p95 The engine Not available: llama.cpp keeps no histograms
Time between tokens p50 (the engine) Mean: seconds spent generating over tokens generated
Request latency, p50 and p95 The engine How long clients waited at the gateways
Generation speed — llama.cpp's own mean, tokens per second while generating
Requests running and waiting Running, waiting Slots busy, requests deferred
Preemptions per second The engine —
KV cache in use The engine, 0–100 % The engine, 0–100 %
GPU utilisation, GPU memory used and total, power, temperature Each replica's GPUs Each replica's GPUs
CPU and memory Each replica Each replica

Gateways (the gateway on your machines and the hosted one):

  • Requests per second by status: 2xx, 4xx (the request was refused, or a limit was reached) and 5xx (the model failed, could not be reached, or did not answer in time).
  • Latency as clients saw it, p50 and p95: from the request arriving to the last byte of the answer (for a streamed answer, the whole stream).
  • Tokens per second by API key. Calls from another workspace under a share show as shared:<their workspace's namespace>, not under their key's name.

Only counts and durations are recorded: never a prompt or an answer.

Cold starts: how many times a request woke the deployment at zero in the last 24 hours (7 days with that range), and how long it took until a replica served — mean and longest.

Read it in the console#

  1. Open Eos → Deployments and pick the deployment.
  2. Open the Metrics tab. It opens Live:
    • a row of the numbers that matter now: tokens per second, requests running and waiting, time to the first token (request latency for llama.cpp), KV cache, GPU utilisation and memory, gateway requests per second, and errors per second (4xx · 5xx);
    • charts that move every 2 seconds over the last 15 minutes, with updated N s ago beside them.
  3. Choose 15 min, 1 hour, 24 hours or 7 days to see the history instead (refreshed every 30 seconds); Live returns to live updates.
  4. All replicas adds the replicas up (rates and memory are summed; utilisation and KV cache averaged; quantiles are over every replica's requests). Per replica draws one line per replica.

A hidden browser tab stops live updates until it is shown again. In live mode, GPU power, temperature and total memory are in the history only: choose a range to see them.

An empty chart says why:

Message Meaning
no replica runs (scaled to zero) The deployment is at zero: nothing runs, so neither the engine nor its machine reports. The next request starts a replica.
the engine has not reported this yet The replica is still loading the model, or the engine does not expose that measure.
no requests through the gateways in this range No program called it through a gateway. The Playground is not counted here.

Watch it from the command line#

astra eos top chat          # redrawn every 2 seconds; Ctrl-C to quit
astra eos top chat --once   # print one reading and exit

The first lines are the deployment's throughput, latency, queue, KV cache and gateway numbers; then one line per replica: its machine, whether it serves, tokens per second, requests running and waiting, KV cache, GPU, GPU memory, and how long ago its machine last reported. A measure the engine does not have is left out.

Read it from the API#

curl -H "Authorization: Bearer $ASTRALYX_TOKEN" \
  "https://api.astralyx.cloud/v1/eos/deployments/chat/metrics?range=1h&per=replica"
Parameter Values Default
range 15m, 1h, 6h, 24h, 7d 1h
per total (the deployment's), replica (one series per replica) total
measures Measure ids, comma-separated (generation_tokens_per_second,ttft_p95) all

Each measure comes with its unit, its source, the series it was read from and its points ([Unix milliseconds, value], about 240 per range). GET …/eos/deployments/chat/metrics/live streams server-sent events: one measures event, then a reading every 2 seconds with each measure's value now, per replica, and each replica's last report. See the Eos API reference.

Get alerted#

Alert rules (Organisation → Alerts) fire on events; see Alerts. Two rules that watch deployments:

  • When a deployment goes down: a rule for deployment_down in the deployment's workspace, sent to your team's Slack channel. It fires when a deployment failed, or had no replica serving for 5 minutes after serving, while not scaled to zero on purpose.
  • When its machines are in trouble: a rule for machine_down and gpu_degraded in every workspace, sent to a webhook. It fires when a machine stops answering, or its GPUs are not healthy (a fault, no driver).

Alert rules do not watch thresholds on these numbers. To be told when the error rate or the latency is too high, read GET …/metrics?range=15m from your own monitoring every minute and compare, for example, the 5xx series of gateway_requests to its total, or the last point of gateway_latency_p95.