Skip to content

Set budgets and model routes#

This page covers the limit on each run, budgets per agent or per workspace per day or month, reading what was spent, and routing some model calls to another model with a fallback. The concepts are in Budgets, model routes and cost.

Before you begin#

  • The limit per run and model routes are part of an agent version: editor or admin. Budgets per period: admin (editors read them).
  • For the API: ASTRA_TOKEN and WS as in the REST API page. The CLI has no budget commands.

Limit each run#

In the agent's New version (or New agent), Limits:

  1. Longest run, minutes: stopped past it (60 by default).
  2. Stop at, per run: dollars of model calls and the provider's tools, at the cluster's prices, checked before each call.
  3. Tokens per run: input and output together.
  4. When reached: Stop the run, or Ask for approval (the run waits for an approval that raises its limit).
spec fragment
{"budget": {"max_seconds": 1800, "max_cost_usd": 2.0, "max_tokens": 400000, "on_exceed": "approval"}}

max_seconds 60 to 31 622 400 (366 days); max_cost_usd and max_tokens more than 0; on_exceed stop (the default) or approval, which needs one of the two limits.

When a run reaches its limit with Stop, its next model call is refused with budget reached: $2.00 of $2.00 (or … tokens), and the agent ends without its model.

Set a budget for an agent or the workspace#

  1. Open Anemoi → Budgets and press New budget.
  2. Name; For: Every agent of the workspace, together or one agent; Per: Day (UTC) or Month (UTC).
  3. Dollars, Tokens, or both.
  4. When reached: Stop, or Ask for approval with Each approval allows, dollars (or tokens).
  5. Press Create. The list shows what was spent in the period and the state, Ok or Exhausted.

Budgets: what was spent, the limits, and the model prices

$ curl -sS -X POST "$WS/agent-budgets" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"metadata": {"name": "daily-agents"}, "spec": {"scope": "workspace", "period": "day", "max_cost_usd": 50}}' \
    | jq '{spec, state: .status.state, observed}'
{
  "spec": {"max_cost_usd": 50.0, "period": "day", "scope": "workspace"},
  "state": "Ok",
  "observed": {"model_calls": 0, "spent_tokens": 0, "spent_usd": 0.0, "stopped_runs": 0}
}
$ curl -sS -X POST "$WS/agent-budgets" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"metadata": {"name": "researcher-monthly"}, "spec": {"scope": "agent:researcher", "period": "month", "max_tokens": 5000000, "on_exceed": "approval"}}'

PUT $WS/agent-budgets/{name} with {spec} changes it; DELETE removes it (and the stops it caused are lifted).

Field Type Default Description
scope string required workspace, or agent:<name>.
period string day day (from 00:00 UTC) or month (from the 1st, 00:00 UTC).
max_cost_usd number — Dollars per period, more than 0.
max_tokens integer — Tokens (input and output) per period, more than 0. At least one of the two.
on_exceed string stop stop, or approval.
grant_usd number a tenth of max_cost_usd With approval: dollars more per approval.
grant_tokens integer a tenth of max_tokens With approval: tokens more per approval.

While a budget is Exhausted, new runs in its scope are refused with 409 BUDGET_EXHAUSTED (budget daily-agents reached: $51.20 of $50.00 this day), and running ones lose their model at their next call (or ask). Its history records each change of state.

See what was spent#

Budgets → Spent shows the workspace's dollars, tokens and model calls Today or This month, by agent and by run; an agent's Costs tab shows its own.

$ curl -sS "$WS/agent-costs?period=month" -H "Authorization: Bearer $ASTRA_TOKEN" | jq '{total, items, budgets: [.budgets[].metadata.name]}'
$ curl -sS "$WS/agents/researcher/costs?period=day" -H "Authorization: Bearer $ASTRA_TOKEN"
$ curl -sS "$WS/model-prices" -H "Authorization: Bearer $ASTRA_TOKEN" | jq '.defaults | length'

total and each item carry model_calls, input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, server_tool_calls and cost_usd.

Route model calls#

Send short requests to a small model and keep the large one for the rest, or fall back to another provider when the first is down.

In Advanced → Model → Routing: Add a rule (its conditions and its model), and tick A fallback when the chosen model fails (5xx, 429, or no answer in time) to choose the fallback. Every model must speak the same API as the primary.

spec fragment
{
  "model": {
    "provider": "openai", "model": "gpt-5", "credential": "openai-key",
    "routes": [
      {"when": {"max_input_tokens": 2000},
       "model": {"provider": "openai", "model": "gpt-5-mini", "credential": "openai-key"}},
      {"when": {"contains_any": ["translate", "summarise"]},
       "model": {"deployment": "chat"}}
    ],
    "fallback": {"provider": "openrouter", "model": "openai/gpt-5", "credential": "openrouter-key"}
  }
}

Every route and the fallback must speak the agent's API: with an Anthropic model they are Anthropic models; with an OpenAI-speaking model, any OpenAI-compatible provider or an Eos deployment. Otherwise the version is refused (… speaks OpenAI's API; the agent's model speaks Anthropic's API).

when field Holds when
max_input_tokens The request is at most this many tokens (characters of its messages over four).
min_input_tokens At least this many.
contains_any Any of these words (at most 32, each at most 64 bytes) is in the user's messages, case ignored.
tool_count_min The request offers at least this many tools.

At most 8 rules, tried in order; a route has no routes or fallback of its own. The trace names the route of each call: primary, route:<n>, fallback.