Skip to content

Evaluations and promotion#

A new version of an agent — a new prompt, model or tool — can make it worse. Evaluations answer is this version good enough? before it takes real work: an eval suite runs the version on test cases and grades the answers; a gate refuses to make a version current until it passes; traffic gives a new version a share of runs first; rollback puts things back in one step.

An agent's Quality tab: the gate, and each version's latest evaluation

Eval suites#

A suite holds:

  • Cases (at most 200): an input, optionally the expected answer, criteria for a judge and tags. A finished run can be turned into a case with Add to eval suite: its input, and its answer as the expected one, read from its machine at that moment.
  • Graders (1 to 10), each grading every case:

    Grader Passes when
    exact The answer is the case's expected (trimmed; ignore_case optional).
    contains The answer contains value (the case's expected when empty).
    regex The answer matches pattern (at most 1 KiB).
    json_path_equals The answer is JSON (or holds a JSON block) whose value at path ($.a.b[0]) is value (the case's expected when absent).
    llm_judge A model scores the answer 0 to 1 by a rubric (and the case's criteria and expected answer); it passes at threshold, 0.5 by default.

    A case's score is the graders' weighted mean (weight, 1 by default); the case passes when every grader passes. - A pass rule: min_score (the mean of the cases' scores, 0.8 by default), optionally min_cases_passed, and no_regression_vs_current (score at least what the current version scored). - The agent it is for (or any agent of the workspace) and a timeout_seconds (60 to 86 400; 2 hours by default).

Eval runs#

Evaluating a version starts an eval run:

  1. Each case runs as a run of that version, at most 4 at a time. These runs are real agent runs — same sandbox, policies, guardrails and budgets (they count toward budgets) — but they are not listed among the agent's runs or its traffic statistics.
  2. Each answer is written to the eval run's drive, on the machine that ran it.
  3. On each machine that ran cases, a small grading run reads the answers there and reports only scores and short notes (at most 500 characters a case). The answers never leave your machines. An llm_judge grader calls its model through that machine's gateway, like an agent.
  4. The verdict — Passed or Failed, with the score, the cases passed and the cost — is recorded for the version.

An eval run is Pending, Running, Grading, then Passed, Failed (also when it timed out) or Cancelled. Its page compares each case with the same suite's latest run on the current version.

A verdict is for the suite as it was: changing a suite's cases, graders or pass rule forgets its verdicts, and versions must be evaluated again.

The gate#

An agent's gate names a suite. While it is required, a version becomes current — by making it current, by writing it as current, or by giving it traffic — only if its latest evaluation by that suite passed (and, with no_regression_vs_current, scored at least the current version's). Otherwise the change is refused with 409 EVAL_GATE_FAILED. The version already current is never blocked.

Promote does it in one step: if the version has passed, it becomes current at once; if not, an eval run starts and makes it current when it passes.

Traffic and rollback#

Traffic shares the runs that name no version among up to 10 versions by weight (the weights sum to 100): a canary such as 90/10. The version is picked from the run's name, so the same name always gets the same version. Every version given a share passes the gate first. The Traffic tab shows each version's runs, success rate, cost and latest evaluation. Making a version current ends the split.

Rollback ends the split (all runs to the current version), or, when there is none, makes the previous current version current again — without the gate, since it was in production. Every change is in the agent's History, with who made it.