Evaluate and promote a new agent version#
You have researcher in production and a better prompt for it. You will
build an eval suite from runs that went well, require it, write version 2
without making it current, evaluate it against the current version, give
it 10 % of runs, then all of them — and know how to undo it in one step.
Before you begin#
- An agent with a few finished runs whose answers you judge good — here
researcher, version 1 current. - The editor role.
- A model to judge with: here the Eos deployment
chat(a provider with a credential works too).
1. Build the suite from real runs#
- Anemoi → Evals → New suite: name
researcher-basics, agentresearcher. Add a grader contains with Valuehttps://(every answer cites a link) and a judge with the Rubric Five bullet points, each with a source URL, all on the question asked, Passes at0.7, Judge modelchat. Minimum score0.8, and require no regression against the current version. Create the suite. - Open each good run of
researcherand press Add to eval suite →researcher-basics, with a criterion or two (names the version). Each becomes a case: its input, and its answer as the expected one. - Add a few hard cases by hand on the suite's Cases and graders tab.
$ curl -sS -X POST "$WS/eval-suites" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" -d '{
"metadata": {"name": "researcher-basics"},
"spec": {"agent": "researcher",
"graders": [{"kind": "contains", "value": "https://"},
{"kind": "llm_judge", "rubric": "Five bullet points, each with a source URL, all on the question asked.",
"threshold": 0.7, "model": {"deployment": "chat"}}],
"pass": {"min_score": 0.8, "no_regression_vs_current": true}}}'
$ for r in researcher-d2b797 researcher-7a2c90 researcher-3f1b55; do
curl -sS -X POST "$WS/eval-suites/researcher-basics/cases/from-run" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d "{\"agent\": \"researcher\", \"run\": \"$r\"}" | jq -r .case.id
done
2. Measure the current version#
On the suite, Evaluate a version → researcher, v1. Wait for
passed or failed; the score is the baseline version 2 must reach.
Each case runs as a real run of v1 (at most 4 at a time), and is graded on the machine that ran it: only scores and short notes come back.
3. Require the gate#
From now on no version becomes current, or takes traffic, without passing.

4. Write version 2, not current#
New version, change the instructions, and save. With the gate required, it is written without becoming current.
$ curl -sS "$WS/agents/researcher" -H "Authorization: Bearer $ASTRA_TOKEN" \
| jq '{spec: (.spec | .instructions += " Prefer primary sources: release notes, specifications, vendor documentation."), current: false}' \
| curl -sS -X PUT "$WS/agents/researcher" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" -d @- \
| jq '{current, latest}'
{"current": 1, "latest": 2}
Written with current left at its default, it is refused with 409
EVAL_GATE_FAILED: version 2 would become current before it is
evaluated by suite researcher-basics….
5. Evaluate it#
Quality → Evaluate next to v2. The eval run's page compares each case with v1's latest run: score, passed, notes.
$ E=$(curl -sS -X POST "$WS/eval-runs" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" \
-d '{"spec": {"suite": "researcher-basics", "version": 2}}' | jq -r .metadata.name)
$ curl -sS "$WS/eval-runs/$E" -H "Authorization: Bearer $ASTRA_TOKEN" \
| jq '{state: .status.state, verdict, verdict_reason, score: .totals.score, baseline: .baseline.score}'
A fail says why: score 0.71 is below 0.80, or below the current version's score.
6. Give it 10 % of runs#
Traffic → Split: v1 90 %, Add a version v2 10 %, Save the split. By version shows runs, success rate, cost per run and the evaluation of each, as real runs come in.
Runs started without a version are shared by their name: the same run name always gets the same version. A version that has not passed cannot be given a share.
7. Promote it, or roll back#
Quality → Promote next to v2: it passed, so it becomes current at once, and the split ends. If v2 misbehaves, Traffic → Roll back to v1 makes v1 current again in one step — without the gate, since it was in production.
$ curl -sS -X POST "$WS/agents/researcher/promote" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{"version": 2}' | jq '{promoted, reason}'
{"promoted": true, "reason": "Version 2 passed suite researcher-basics (score 0.92) and is now current"}
$ curl -sS -X POST "$WS/agents/researcher/rollback" -H "Authorization: Bearer $ASTRA_TOKEN" \
-H "Content-Type: application/json" -d '{}'
Promote on a version that has not passed starts an eval run that makes it current when it passes. Every change — gate, split, promotion, rollback — is in the agent's History with who made it, and in evidence packs.
Keep the suite honest#
- Changing the suite's cases, graders or pass rule forgets its verdicts: versions are evaluated again before the gate lets them through.
- Evaluation runs count toward budgets like any run; the eval run's cost is on its page.