Run evaluations and promote a version#
This page covers eval suites, eval runs, the gate, promotion, traffic splits and rollback. The concepts are in Evaluations and promotion.
Before you begin#
- The editor or admin role (viewers read suites, eval runs and an agent's quality).
- For the API:
ASTRA_TOKENandWSas in the REST API page. The CLI has no evaluation commands.
Write a suite#
- Open Anemoi → Evals and press New suite.
- Name, the Agent it is for (or any agent of the workspace), a Description.
- Cases: for each, an Id, the Input, the Expected answer if there is one, Criteria for a judge, Tags; or Import them as JSON (Cases as JSON).
- Graders: Add a grader — exact, contains (Value), regex (Pattern), JSON path equals (JSON path, Value (JSON or text)), judge (Rubric, Passes at, Judge model) — each with a Weight.
- Passing: Minimum score (0.8 by default), Cases that must pass, and whether a version must score no less than the current one.
- Press Create the suite.

{
"metadata": {"name": "researcher-basics"},
"spec": {
"agent": "researcher",
"description": "Answers cite their sources and stay on the question.",
"cases": [
{"id": "rust-release", "input": "What changed in the latest stable Rust release?",
"criteria": ["names the version", "every bullet has a URL"]},
{"id": "cuda-13", "input": "Which GPUs did CUDA 13 drop?", "expected": "Maxwell",
"criteria": ["names the architectures dropped"]}
],
"graders": [
{"kind": "contains", "value": "https://"},
{"kind": "llm_judge", "rubric": "Five bullet points, each with a source URL, all on the question asked.",
"threshold": 0.7, "model": {"deployment": "chat"}}
],
"pass": {"min_score": 0.8, "no_regression_vs_current": true}
}
}
$ curl -sS -X POST "$WS/eval-suites" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d @researcher-basics.json | jq -r .digest
sha256:a1fcf1bd165ed1ef702600ca426a36d659f2d368ec514306f08d374fde886fbe
PUT $WS/eval-suites/{name} with {spec} replaces it; GET
$WS/eval-suites?agent=researcher lists the suites that may evaluate an
agent; DELETE removes one (refused with 409 EVAL_SUITE_IN_USE
while a gate names it). See the fields.
Every judge of a suite uses the same model; its calls go through the grading machine's gateway and are priced like the agent's.
Turn real runs into cases#
On a finished run's page press Add to eval suite, choose the suite, optionally criteria and tags, and confirm: the run's input becomes a case, and its answer — read from its machine now — the expected one.
$ curl -sS -X POST "$WS/eval-suites/researcher-basics/cases/from-run" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" \
-d '{"agent": "researcher", "run": "researcher-d2b797", "criteria": ["cites the release notes"]}'
id (from the run's name when absent), tags, and
expected_from_answer (true by default). 409 RUN_ANSWER_UNAVAILABLE
when the run's machine cannot give its answer (offline, or the answer
gone).
A case keeps where it came from (source: agent, run, version, who, when).
Evaluate a version#
On the suite's page press Evaluate a version, choose the agent and version. Or, on the agent's Quality tab, Evaluate next to a version. The eval run's page shows each case — its run, score, passed, the graders' notes — compared with the current version's latest run, and the verdict.

$ curl -sS -X POST "$WS/eval-runs" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{"spec": {"suite": "researcher-basics", "version": 2}}' \
| jq -r .metadata.name
researcher-basics-4e45de
$ curl -sS "$WS/eval-runs/researcher-basics-4e45de" -H "Authorization: Bearer $ASTRA_TOKEN" \
| jq '{state: .status.state, totals, verdict, verdict_reason, cases: [.cases[] | {id, state, score, passed, notes}]}'
spec: suite, agent (the suite's when absent), version (the
current one when absent), promote. GET $WS/eval-runs?suite=&agent=&version=&state=
lists them; POST …/cancel stops one; DELETE removes one that ended.
The eval run goes Pending → Running (cases at most 4 at a time) →
Grading → Passed or Failed. It fails as a whole when it exceeds the
suite's timeout_seconds.
Require the gate#
On the agent's Quality tab, under Gate, choose the Suite, tick Required: a version that has not passed cannot be made current and press Save the gate. Versions shows each version's latest evaluation per suite.
$ curl -sS -X PUT "$WS/agents/researcher/gate" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{"suite": "researcher-basics", "required": true}'
$ curl -sS "$WS/agents/researcher/quality" -H "Authorization: Bearer $ASTRA_TOKEN" | jq '{current, gate, verdicts}'
DELETE $WS/agents/researcher/gate removes it. The suite must be of
the workspace, and for this agent or any.
From now on, writing a new version as current, making a version current or giving it traffic is refused unless it passed:
{"code": "EVAL_GATE_FAILED", "message": "version 2 has not been evaluated by suite researcher-basics (the gate requires it): evaluate it, or POST /agents/researcher/promote with it"}
Promote#
On the Quality tab, Promote next to the version.
$ curl -sS -X POST "$WS/agents/researcher/promote" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{"version": 2}' | jq '{promoted, reason, eval_run: .eval_run.metadata.name}'
{
"promoted": false,
"reason": "Evaluating version 2 with suite researcher-basics: it becomes current when it passes",
"eval_run": "researcher-basics-4e45de"
}
200 with promoted: true when the version had passed already (or
there is no gate and no suite given); 202 with the eval run it
started otherwise. suite evaluates by another suite than the gate's.
Split traffic and roll back#
On the agent's Traffic tab, under Split, give versions their share (Add a version, the slider or the percentage) and press Save the split. By version shows each version's runs, succeeded, failed, success rate, cost, cost per run and latest evaluation. Roll back to v… ends the split (all runs to the current version), or, with no split, makes the previous version current again.

$ curl -sS -X PUT "$WS/agents/researcher/traffic" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{"traffic": [{"version": 1, "weight": 90}, {"version": 2, "weight": 10}]}'
$ curl -sS "$WS/agents/researcher/traffic" -H "Authorization: Bearer $ASTRA_TOKEN" | jq '.versions[] | {version, weight, runs, success_rate, cost_usd}'
$ curl -sS -X POST "$WS/agents/researcher/rollback" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{}'
Weights sum to 100, at most 10 versions, each passing the gate; []
sends everything to the current version. rollback takes an optional
version, which must be the previous one.
A split applies to runs that name no version; the version is picked from the run's name, so the same name always gets the same version.
Suite fields#
| Field | Type | Default | Description |
|---|---|---|---|
agent |
string | any agent | The agent it is for. |
description |
string | — | At most 4 000 bytes. |
cases[].id |
string | case-<n> |
[a-z0-9][a-z0-9_-]*, at most 40, unique. |
cases[].input |
string | required | At most 32 KiB. |
cases[].expected |
string | — | At most 32 KiB. |
cases[].criteria |
strings | — | At most 20, each at most 1 000 bytes. |
cases[].tags |
strings | — | At most 10, each 1 to 40 bytes. |
graders[].kind |
string | required | exact, contains, regex, json_path_equals, llm_judge. |
graders[].name |
string | its kind | Unique in the suite; at most 40. |
graders[].weight |
number | 1 | Its weight in a case's score. |
graders[].ignore_case |
bool | false |
exact, contains. |
graders[].value |
string / JSON | the case's expected |
contains, json_path_equals. |
graders[].pattern |
string | required | regex (Rust syntax), at most 1 KiB. |
graders[].path |
string | required | json_path_equals: $.a.b[0]. |
graders[].rubric |
string | required | llm_judge. |
graders[].threshold |
number | 0.5 | llm_judge: the score it passes at. |
graders[].model |
object | required | llm_judge: a deployment, or a provider with a credential (as an agent's model, no routes). |
pass.min_score |
number | 0.8 | 0 to 1. |
pass.min_cases_passed |
integer | — | At most 200. |
pass.no_regression_vs_current |
bool | false |
Score at least the current version's. |
timeout_seconds |
integer | 7200 | 60 to 86 400. |
At most 200 cases and 10 graders; the suite at most 512 KiB. Answers longer than 1 MiB are graded on their first 1 MiB.