Skip to content

Run evaluations and promote a version#

This page covers eval suites, eval runs, the gate, promotion, traffic splits and rollback. The concepts are in Evaluations and promotion.

Before you begin#

  • The editor or admin role (viewers read suites, eval runs and an agent's quality).
  • For the API: ASTRA_TOKEN and WS as in the REST API page. The CLI has no evaluation commands.

Write a suite#

  1. Open Anemoi → Evals and press New suite.
  2. Name, the Agent it is for (or any agent of the workspace), a Description.
  3. Cases: for each, an Id, the Input, the Expected answer if there is one, Criteria for a judge, Tags; or Import them as JSON (Cases as JSON).
  4. Graders: Add a grader — exact, contains (Value), regex (Pattern), JSON path equals (JSON path, Value (JSON or text)), judge (Rubric, Passes at, Judge model) — each with a Weight.
  5. Passing: Minimum score (0.8 by default), Cases that must pass, and whether a version must score no less than the current one.
  6. Press Create the suite.

A suite's cases and graders

researcher-basics.json
{
  "metadata": {"name": "researcher-basics"},
  "spec": {
    "agent": "researcher",
    "description": "Answers cite their sources and stay on the question.",
    "cases": [
      {"id": "rust-release", "input": "What changed in the latest stable Rust release?",
       "criteria": ["names the version", "every bullet has a URL"]},
      {"id": "cuda-13", "input": "Which GPUs did CUDA 13 drop?", "expected": "Maxwell",
       "criteria": ["names the architectures dropped"]}
    ],
    "graders": [
      {"kind": "contains", "value": "https://"},
      {"kind": "llm_judge", "rubric": "Five bullet points, each with a source URL, all on the question asked.",
       "threshold": 0.7, "model": {"deployment": "chat"}}
    ],
    "pass": {"min_score": 0.8, "no_regression_vs_current": true}
  }
}
$ curl -sS -X POST "$WS/eval-suites" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d @researcher-basics.json | jq -r .digest
sha256:a1fcf1bd165ed1ef702600ca426a36d659f2d368ec514306f08d374fde886fbe

PUT $WS/eval-suites/{name} with {spec} replaces it; GET $WS/eval-suites?agent=researcher lists the suites that may evaluate an agent; DELETE removes one (refused with 409 EVAL_SUITE_IN_USE while a gate names it). See the fields.

Every judge of a suite uses the same model; its calls go through the grading machine's gateway and are priced like the agent's.

Turn real runs into cases#

On a finished run's page press Add to eval suite, choose the suite, optionally criteria and tags, and confirm: the run's input becomes a case, and its answer — read from its machine now — the expected one.

$ curl -sS -X POST "$WS/eval-suites/researcher-basics/cases/from-run" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"agent": "researcher", "run": "researcher-d2b797", "criteria": ["cites the release notes"]}'

id (from the run's name when absent), tags, and expected_from_answer (true by default). 409 RUN_ANSWER_UNAVAILABLE when the run's machine cannot give its answer (offline, or the answer gone).

A case keeps where it came from (source: agent, run, version, who, when).

Evaluate a version#

On the suite's page press Evaluate a version, choose the agent and version. Or, on the agent's Quality tab, Evaluate next to a version. The eval run's page shows each case — its run, score, passed, the graders' notes — compared with the current version's latest run, and the verdict.

An eval suite and its evaluations

$ curl -sS -X POST "$WS/eval-runs" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"spec": {"suite": "researcher-basics", "version": 2}}' \
    | jq -r .metadata.name
researcher-basics-4e45de
$ curl -sS "$WS/eval-runs/researcher-basics-4e45de" -H "Authorization: Bearer $ASTRA_TOKEN" \
    | jq '{state: .status.state, totals, verdict, verdict_reason, cases: [.cases[] | {id, state, score, passed, notes}]}'

spec: suite, agent (the suite's when absent), version (the current one when absent), promote. GET $WS/eval-runs?suite=&agent=&version=&state= lists them; POST …/cancel stops one; DELETE removes one that ended.

The eval run goes Pending → Running (cases at most 4 at a time) → Grading → Passed or Failed. It fails as a whole when it exceeds the suite's timeout_seconds.

Require the gate#

On the agent's Quality tab, under Gate, choose the Suite, tick Required: a version that has not passed cannot be made current and press Save the gate. Versions shows each version's latest evaluation per suite.

$ curl -sS -X PUT "$WS/agents/researcher/gate" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"suite": "researcher-basics", "required": true}'
$ curl -sS "$WS/agents/researcher/quality" -H "Authorization: Bearer $ASTRA_TOKEN" | jq '{current, gate, verdicts}'

DELETE $WS/agents/researcher/gate removes it. The suite must be of the workspace, and for this agent or any.

From now on, writing a new version as current, making a version current or giving it traffic is refused unless it passed:

{"code": "EVAL_GATE_FAILED", "message": "version 2 has not been evaluated by suite researcher-basics (the gate requires it): evaluate it, or POST /agents/researcher/promote with it"}

Promote#

On the Quality tab, Promote next to the version.

$ curl -sS -X POST "$WS/agents/researcher/promote" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"version": 2}' | jq '{promoted, reason, eval_run: .eval_run.metadata.name}'
{
  "promoted": false,
  "reason": "Evaluating version 2 with suite researcher-basics: it becomes current when it passes",
  "eval_run": "researcher-basics-4e45de"
}

200 with promoted: true when the version had passed already (or there is no gate and no suite given); 202 with the eval run it started otherwise. suite evaluates by another suite than the gate's.

Split traffic and roll back#

On the agent's Traffic tab, under Split, give versions their share (Add a version, the slider or the percentage) and press Save the split. By version shows each version's runs, succeeded, failed, success rate, cost, cost per run and latest evaluation. Roll back to v… ends the split (all runs to the current version), or, with no split, makes the previous version current again.

An agent's Traffic tab

$ curl -sS -X PUT "$WS/agents/researcher/traffic" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"traffic": [{"version": 1, "weight": 90}, {"version": 2, "weight": 10}]}'
$ curl -sS "$WS/agents/researcher/traffic" -H "Authorization: Bearer $ASTRA_TOKEN" | jq '.versions[] | {version, weight, runs, success_rate, cost_usd}'
$ curl -sS -X POST "$WS/agents/researcher/rollback" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{}'

Weights sum to 100, at most 10 versions, each passing the gate; [] sends everything to the current version. rollback takes an optional version, which must be the previous one.

A split applies to runs that name no version; the version is picked from the run's name, so the same name always gets the same version.

Suite fields#

Field Type Default Description
agent string any agent The agent it is for.
description string — At most 4 000 bytes.
cases[].id string case-<n> [a-z0-9][a-z0-9_-]*, at most 40, unique.
cases[].input string required At most 32 KiB.
cases[].expected string — At most 32 KiB.
cases[].criteria strings — At most 20, each at most 1 000 bytes.
cases[].tags strings — At most 10, each 1 to 40 bytes.
graders[].kind string required exact, contains, regex, json_path_equals, llm_judge.
graders[].name string its kind Unique in the suite; at most 40.
graders[].weight number 1 Its weight in a case's score.
graders[].ignore_case bool false exact, contains.
graders[].value string / JSON the case's expected contains, json_path_equals.
graders[].pattern string required regex (Rust syntax), at most 1 KiB.
graders[].path string required json_path_equals: $.a.b[0].
graders[].rubric string required llm_judge.
graders[].threshold number 0.5 llm_judge: the score it passes at.
graders[].model object required llm_judge: a deployment, or a provider with a credential (as an agent's model, no routes).
pass.min_score number 0.8 0 to 1.
pass.min_cases_passed integer — At most 200.
pass.no_regression_vs_current bool false Score at least the current version's.
timeout_seconds integer 7200 60 to 86 400.

At most 200 cases and 10 graders; the suite at most 512 KiB. Answers longer than 1 MiB are graded on their first 1 MiB.