Evaluations and promotion#
A new version of an agent — a new prompt, model or tool — can make it worse. Evaluations answer is this version good enough? before it takes real work: an eval suite runs the version on test cases and grades the answers; a gate refuses to make a version current until it passes; traffic gives a new version a share of runs first; rollback puts things back in one step.

Eval suites#
A suite holds:
- Cases (at most 200): an
input, optionally theexpectedanswer,criteriafor a judge andtags. A finished run can be turned into a case with Add to eval suite: its input, and its answer as the expected one, read from its machine at that moment. -
Graders (1 to 10), each grading every case:
Grader Passes when exactThe answer is the case's expected(trimmed;ignore_caseoptional).containsThe answer contains value(the case'sexpectedwhen empty).regexThe answer matches pattern(at most 1 KiB).json_path_equalsThe answer is JSON (or holds a JSON block) whose value at path($.a.b[0]) isvalue(the case'sexpectedwhen absent).llm_judgeA model scores the answer 0 to 1 by a rubric(and the case's criteria and expected answer); it passes atthreshold, 0.5 by default.A case's score is the graders' weighted mean (
weight, 1 by default); the case passes when every grader passes. - A pass rule:min_score(the mean of the cases' scores, 0.8 by default), optionallymin_cases_passed, andno_regression_vs_current(score at least what the current version scored). - The agent it is for (or any agent of the workspace) and atimeout_seconds(60 to 86 400; 2 hours by default).
Eval runs#
Evaluating a version starts an eval run:
- Each case runs as a run of that version, at most 4 at a time. These runs are real agent runs — same sandbox, policies, guardrails and budgets (they count toward budgets) — but they are not listed among the agent's runs or its traffic statistics.
- Each answer is written to the eval run's drive, on the machine that ran it.
- On each machine that ran cases, a small grading run reads the answers
there and reports only scores and short notes (at most 500
characters a case). The answers never leave your machines. An
llm_judgegrader calls its model through that machine's gateway, like an agent. - The verdict —
PassedorFailed, with the score, the cases passed and the cost — is recorded for the version.
An eval run is Pending, Running, Grading, then Passed, Failed
(also when it timed out) or Cancelled. Its page compares each case with
the same suite's latest run on the current version.
A verdict is for the suite as it was: changing a suite's cases, graders or pass rule forgets its verdicts, and versions must be evaluated again.
The gate#
An agent's gate names a suite. While it is required, a version
becomes current — by making it current, by writing it as current, or by
giving it traffic — only if its latest evaluation by that suite passed
(and, with no_regression_vs_current, scored at least the current
version's). Otherwise the change is refused with 409 EVAL_GATE_FAILED.
The version already current is never blocked.
Promote does it in one step: if the version has passed, it becomes current at once; if not, an eval run starts and makes it current when it passes.
Traffic and rollback#
Traffic shares the runs that name no version among up to 10 versions by weight (the weights sum to 100): a canary such as 90/10. The version is picked from the run's name, so the same name always gets the same version. Every version given a share passes the gate first. The Traffic tab shows each version's runs, success rate, cost and latest evaluation. Making a version current ends the split.
Rollback ends the split (all runs to the current version), or, when there is none, makes the previous current version current again — without the gate, since it was in production. Every change is in the agent's History, with who made it.