Budgets, model routes and cost#
Every model call an agent run makes goes through the gateway on its machine, which counts its tokens and prices it. Limits stop (or ask before going on) at two levels: each run, and each agent or the whole workspace per day or month. Routing rules send some calls to another model, and a fallback takes a call when the model fails. This page explains each.
What is counted#
For every model call the gateway records the model, the route that took it, input tokens (cached ones included), output tokens, tokens read from and written to the provider's cache, uses of the provider's own tools (web search…), and the cost in US dollars. A run's totals are its spend, shown on its page, in its agent's Costs tab, in Budgets, and in its receipt.
Prices. A call is priced per million tokens (input, output, cache read,
cache write) and per use of a provider's tool, from the cluster's price
table: the vendors' public list prices for Anthropic, OpenAI, DeepSeek,
Groq, Mistral and OpenRouter models, under any prices set for the cluster.
The most specific entry for the model wins (claude-sonnet-4-5* over
claude-*). A run takes its price when it is made: a price change
applies to new runs. An Eos deployment costs nothing per token unless
priced for the cluster (its GPU hours are already in
usage); a model no entry matches costs 0 — its
tokens are still counted, and token limits still hold.
Budgets → Model prices shows the table. Workspaces read it; they
cannot change it (GET $WS/model-prices).
The limit per run#
Set on the agent (each version), checked by the gateway before each model call — exact to one call:
| Field | Console | Meaning |
|---|---|---|
budget.max_seconds |
Longest run, minutes | The run's time limit: 60 s to 366 days; 1 hour when unset. |
budget.max_cost_usd |
Stop at, per run | Dollars of model calls and provider tools, at the run's prices. |
budget.max_tokens |
Tokens per run | Input and output tokens together. |
budget.on_exceed |
When reached | stop (Stop the run): its model calls are refused with budget reached, and the agent ends without its model. approval (Ask for approval): the call waits for an approval of kind budget; approved, the limit is raised. |
Use the per-run limit for hard ceilings.
Budgets per period#
A budget limits what one agent (agent:<name>), or every agent of the
workspace together (workspace), spends per UTC day or calendar
month, in dollars, tokens or both.
While a budget is Exhausted:
- the next model call of each running run in its scope is refused — or,
with
on_exceed: approval, asks a person, and each approval allowsgrant_usdorgrant_tokensmore (a tenth of the limit by default); - new runs in its scope are refused with
409 BUDGET_EXHAUSTED.
It turns Ok again when the period turns over, or when an admin raises or
deletes it.
Period budgets are close, not exact
Spend is reported by machines every few seconds and summed every 15 seconds, so a scope may overspend by what its running runs spend in about 20 seconds, plus the calls they had in flight. For a hard ceiling, set the limit per run.
Workspace admins write budgets; editors and viewers read them. Evaluation runs count toward budgets like any run.
Model routes#
A version's model may carry routing rules (at most 8) and a fallback, decided per call by the gateway from the request alone:
- Rules are tried in order; the first whose conditions all hold picks
its model; none matching leaves the agent's own model. Conditions:
max_input_tokens,min_input_tokens(estimated as the characters of the messages over four),contains_any(words in the user's messages, case ignored, at most 32),tool_count_min(the request offers at least that many tools). - The fallback takes a call when the model it went to answers
429or5xx, or cannot be reached. The call is sent to it once.
Every route speaks the agent's API (Anthropic's Messages API, or OpenAI's),
is a deployment or a provider with its own credential, and is priced on its
own. The trace names the route that took each call: primary, route:<n>
or fallback. A route whose deployment is not ready when the run is made
is left out, and the run goes on with the others.