Skip to content

Evaluations

This group covers evaluating a workflow: this overview, then datasets of cases, experiments that run them, and the typed Result they are matched against.

An evaluation answers one question before you switch a workflow to a new version: does it still produce the right Result, and what does it cost? You keep a dataset of cases, run a workflow version over it as an experiment, and compare that experiment with a baseline.

Evaluations are warn-only. Nothing in them blocks a push or a deploy; you decide what a regression means for your release.

The parts

Term What it is
Dataset A versioned set of cases for one workflow. Names are unique per workflow.
Case One workflow.start payload (input) and what the run’s Result should be (expected).
Dataset version A frozen cut of a dataset: every live case pinned at one revision. Cut automatically when an experiment runs changed cases.
Experiment One dataset version run against one workflow version, each case run repeats times.
Run One case × one repeat inside an experiment, with its output, scores, and metrics.
Result The fields a workflow’s Result step (output step) declares. Runs are scored against these fields and nothing else.
Baseline The experiment other experiments on a dataset are compared to.
dataset "triage" v3 ──┐
├─► experiment exp-… ──► runs (case × repeat) ──► scores + metrics
workflow v1.4.0 ──────┘ │
▼
compare with the baseline experiment

How a run is scored

  1. The experiment publishes the case’s input as a workflow.start event into a context of its own (eval-<experimentId>-<caseId>-<repeat>-a<attempt>).
  2. The workflow runs as it would for any other start event: real model calls, real agent steps, real cubby writes.
  3. When the run ends, its workflow.completed output (the declared Result) is matched field by field against the case’s expected.
  4. The run’s duration, tokens, cost, and step count are read back from the platform and checked against the dataset’s limits.

A case passes when every judged field passes. Limits only warn; they never fail a case. See Results for the matching rules.

An experiment runs your workflow for real. Every run calls models, writes cubby rows, publishes events, and performs connector actions. Run experiments in a vault where that is safe, and make side-effecting steps idempotent. An experiment’s runs have ids starting with run-eval-, so rows you key on {{ $runId }} are easy to tell apart.

Evaluations and tests

Local tests Evaluations
Runs On your machine, in-process On the platform, against a deployed version
Models Mocked Real
Checks Exact behaviour: cubby rows, events, branches, budgets The Result, plus duration, tokens, cost, steps
Answers “Does the graph do what I wrote?” “Is this version as good as the last one, on real inputs?”
Use Every commit Before switching the live version, after prompt or model changes

Use both. Tests pin down logic that must never change; a dataset catches the quality drift that only real model calls show. See Test a workflow.

Where evaluations live

ROC. Open a workflow and select the Evaluations tab. It has two views:

  • Experiments: the headline metrics (Accuracy, p50 duration, Tokens/run, Cost/run, Regressions), a History chart, Accuracy by field, and the experiment list with a Verdict per experiment.
  • Dataset: the cases, Add from executions, Add case, and the dataset’s History of versions.

Run experiment starts an experiment from either view. The Evaluations item in the sidebar lists every dataset across the Agent Service’s workflows, with each one’s Latest verdict and Accuracy.

Experiments view with metrics, history chart and experiment list CLI. cef eval reads and writes the same storage as ROC, so either sees what the other wrote:

Terminal window
cef eval create triage
cef eval pull triage # cases into ./datasets/triage/
cef eval push triage # your edits back as new revisions
cef eval run triage # run against the live version
cef eval compare <base> <candidate>

Library. @cef-ai/eval is the browser-safe library both ROC and the CLI use. Use it to build your own tooling over the same storage.

Storage

Datasets and experiments are JSON objects in your Agent Service’s bucket, under workflows/<workflowId>/. Case revisions and dataset versions are write-once, so an old experiment always reloads exactly the cases it ran. See Datasets.