Skip to content

Watch usage#

Eos counts every answer: prompt tokens, completion tokens and requests. Replicas also hold GPUs (or CPU cores), which count as GPU-hours like any run. This page shows where to see both, how to price them, and how to be told when a deployment goes down.

What is counted#

What Counted Shown under
Tokens and requests through a gateway Per deployment, API key and minute; reported within seconds The deployment's model
Tokens in the Playground The same, under the person who chatted The deployment's model
Calls to a deployment shared with you In your usage too shared:<owner namespace>.<model>
GPU-hours, core-hours, memory While a replica holds them, as for any run The GPU model (NVIDIA H100 80GB HBM3…)

A request counts even when it fails. Tokens count when the engine reports them: for a streamed answer the gateway asks the engine for its usage itself. A stopped Playground answer counts what was written. A deployment scaled to zero holds nothing.

Per key, the console shows only Last used (to the minute); token totals are per model.

See it#

Open Organisation → Usage (organisation members). Choose This month, Last month or Last 30 days.

  • The Tokens tile: tokens in and out over the period.
  • The workspace table: Tokens in / out per workspace, next to GPU-hours, core-hours, GiB-hours, worker-hours and cost.
  • Tokens by model: what deployments answered through a gateway, with an API key — Model, Requests, Prompt tokens, Completion tokens and Token cost.

The organisation's usage page

The Playground's side panel shows the tokens of the current conversation.

$ curl -fsS "https://console.astralyx.cloud/api/v1/orgs/acme/usage?start=2026-10-01T00:00:00Z&end=2026-11-01T00:00:00Z" \
    -H "Authorization: Bearer $ASTRAEUS_TOKEN" | jq '.items[] | {workspace, tokens: .usage.tokens, token_cost: .usage.token_cost}'

Each item's usage.tokens maps a model to {prompt, completion, requests, cost}. See Usage, cost and budgets.

Price tokens#

Each cluster has a price table: on a cluster dedicated to your organisation, its owners and admins set it under Organisation → Clusters & machines → the cluster → Prices; on Astraeus Cloud, Astralyx sets it and you see it. For tokens:

Key Unit
tokens_prompt:<model> Per million prompt tokens of that model (the workspace's model name)
tokens_completion:<model> Per million completion tokens of that model
tokens_prompt, tokens_completion Any model without its own price

A shared model is priced under its own name, at the owner's cluster's prices. Pricing both a deployment's GPU-hours and its tokens charges both; a price table usually prices one or the other.

Alert when a deployment goes down#

The alert deployment_down fires when a deployment failed, or had no replica serving for 5 minutes after it was serving — unless it scaled to zero on purpose. Add a rule for it under Organisation → Alerts; see Alerts.