Skip to content

Evaluate and promote a new agent version#

You have researcher in production and a better prompt for it. You will build an eval suite from runs that went well, require it, write version 2 without making it current, evaluate it against the current version, give it 10 % of runs, then all of them — and know how to undo it in one step.

Before you begin#

  • An agent with a few finished runs whose answers you judge good — here researcher, version 1 current.
  • The editor role.
  • A model to judge with: here the Eos deployment chat (a provider with a credential works too).

1. Build the suite from real runs#

  1. Anemoi → Evals → New suite: name researcher-basics, agent researcher. Add a grader contains with Value https:// (every answer cites a link) and a judge with the Rubric Five bullet points, each with a source URL, all on the question asked, Passes at 0.7, Judge model chat. Minimum score 0.8, and require no regression against the current version. Create the suite.
  2. Open each good run of researcher and press Add to eval suite → researcher-basics, with a criterion or two (names the version). Each becomes a case: its input, and its answer as the expected one.
  3. Add a few hard cases by hand on the suite's Cases and graders tab.
$ curl -sS -X POST "$WS/eval-suites" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" -d '{
    "metadata": {"name": "researcher-basics"},
    "spec": {"agent": "researcher",
      "graders": [{"kind": "contains", "value": "https://"},
                  {"kind": "llm_judge", "rubric": "Five bullet points, each with a source URL, all on the question asked.",
                   "threshold": 0.7, "model": {"deployment": "chat"}}],
      "pass": {"min_score": 0.8, "no_regression_vs_current": true}}}'
$ for r in researcher-d2b797 researcher-7a2c90 researcher-3f1b55; do
    curl -sS -X POST "$WS/eval-suites/researcher-basics/cases/from-run" -H "Authorization: Bearer $ASTRA_TOKEN" \
      -H "Content-Type: application/json" -d "{\"agent\": \"researcher\", \"run\": \"$r\"}" | jq -r .case.id
  done

2. Measure the current version#

On the suite, Evaluate a version → researcher, v1. Wait for passed or failed; the score is the baseline version 2 must reach.

$ curl -sS -X POST "$WS/eval-runs" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"spec": {"suite": "researcher-basics", "version": 1}}' | jq -r .metadata.name

Each case runs as a real run of v1 (at most 4 at a time), and is graded on the machine that ran it: only scores and short notes come back.

3. Require the gate#

researcher → Quality → Gate: suite researcher-basics, tick Required, Save the gate.

$ curl -sS -X PUT "$WS/agents/researcher/gate" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"suite": "researcher-basics", "required": true}'

From now on no version becomes current, or takes traffic, without passing.

The Quality tab: the gate and each version's latest evaluation

4. Write version 2, not current#

New version, change the instructions, and save. With the gate required, it is written without becoming current.

$ curl -sS "$WS/agents/researcher" -H "Authorization: Bearer $ASTRA_TOKEN" \
    | jq '{spec: (.spec | .instructions += " Prefer primary sources: release notes, specifications, vendor documentation."), current: false}' \
    | curl -sS -X PUT "$WS/agents/researcher" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" -d @- \
    | jq '{current, latest}'
{"current": 1, "latest": 2}

Written with current left at its default, it is refused with 409 EVAL_GATE_FAILED: version 2 would become current before it is evaluated by suite researcher-basics….

5. Evaluate it#

Quality → Evaluate next to v2. The eval run's page compares each case with v1's latest run: score, passed, notes.

$ E=$(curl -sS -X POST "$WS/eval-runs" -H "Authorization: Bearer $ASTRA_TOKEN" -H "Content-Type: application/json" \
      -d '{"spec": {"suite": "researcher-basics", "version": 2}}' | jq -r .metadata.name)
$ curl -sS "$WS/eval-runs/$E" -H "Authorization: Bearer $ASTRA_TOKEN" \
    | jq '{state: .status.state, verdict, verdict_reason, score: .totals.score, baseline: .baseline.score}'

A fail says why: score 0.71 is below 0.80, or below the current version's score.

6. Give it 10 % of runs#

Traffic → Split: v1 90 %, Add a version v2 10 %, Save the split. By version shows runs, success rate, cost per run and the evaluation of each, as real runs come in.

$ curl -sS -X PUT "$WS/agents/researcher/traffic" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"traffic": [{"version": 1, "weight": 90}, {"version": 2, "weight": 10}]}'

Runs started without a version are shared by their name: the same run name always gets the same version. A version that has not passed cannot be given a share.

7. Promote it, or roll back#

Quality → Promote next to v2: it passed, so it becomes current at once, and the split ends. If v2 misbehaves, Traffic → Roll back to v1 makes v1 current again in one step — without the gate, since it was in production.

$ curl -sS -X POST "$WS/agents/researcher/promote" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{"version": 2}' | jq '{promoted, reason}'
{"promoted": true, "reason": "Version 2 passed suite researcher-basics (score 0.92) and is now current"}
$ curl -sS -X POST "$WS/agents/researcher/rollback" -H "Authorization: Bearer $ASTRA_TOKEN" \
    -H "Content-Type: application/json" -d '{}'

Promote on a version that has not passed starts an eval run that makes it current when it passes. Every change — gate, split, promotion, rollback — is in the agent's History with who made it, and in evidence packs.

Keep the suite honest#

  • Changing the suite's cases, graders or pass rule forgets its verdicts: versions are evaluated again before the gate lets them through.
  • Evaluation runs count toward budgets like any run; the eval run's cost is on its page.