Skip to content

Datasets

A dataset is the set of cases one workflow is judged on. Each case is a start payload and the Result that payload should produce.

Create a dataset

In ROC. Open the workflow, select Evaluations, then New dataset (or Create empty dataset on a workflow with none). Give it a Name, an optional Description, and optional limits: Max duration (s), Max tokens, Max steps, Max cost ($). The sidebar Evaluations page has the same New dataset button with a Workflow picker.

With the CLI. Run it in the workflow’s project folder; --workflow defaults to the workflow your cef.config.ts (or dist/) declares.

Terminal window
cef eval create triage --description "support tickets by category" --max-duration-ms 30000
Flag Meaning
--description <text> What the dataset covers
--max-duration-ms <n> Warn when a run takes longer
--max-tokens <n> Warn when a run uses more tokens
--max-steps <n> Warn when a run takes more steps
--max-cost <n> Warn when a run costs more

Dataset names match ^[a-z0-9][a-z0-9-]{0,62}$ and are unique per workflow. Every cef eval command needs store credentials.

The Dataset view: cases with their inputs and expected results

Case shape

{
"id": "double-charge",
"input": { "ticketId": "T-1001", "text": "I was charged twice for March" },
"expected": {
"category": "billing",
"confidence": ">= 0.8",
"summary": { "$exists": true }
},
"limits": { "maxDurationMs": 20000 },
"tags": ["billing"],
"notes": "Seen in production twice in one week."
}
Field Required Meaning
id yes [A-Za-z0-9._-], 1–80 characters. Stays taken after the case is deleted.
input yes The workflow.start payload, as an object.
expected no Keyed by Result field; dot paths reach nested fields. A case with no expectations runs but is not scored. See Results.
limits no maxDurationMs, maxTokens, maxSteps, maxCost. Each overrides the dataset’s own.
tags no Free labels for filtering.
notes no Free text.
source no Set when the case was captured from a run: { context, workflowVersion?, capturedAt }.

input is exactly what an experiment publishes as workflow.start, so it must contain everything your trigger step reads.

Add cases by hand

In ROC, on the Dataset view select Add case. Fill in Case id, Input (workflow.start payload), the Expected result, Tags, and Notes. Editing a case and choosing Save revision writes a new revision; the old one stays.

In your repo, pull the dataset, edit the JSON, and push it back:

Terminal window
cef eval pull triage # writes datasets/triage/
$EDITOR datasets/triage/cases/double-charge.json
cef eval push triage

pull writes this layout (change it with --dir):

datasets/triage/header.json the dataset header (without the baseline)
datasets/triage/cases/<caseId>.json each live case at its latest revision
datasets/triage/.versions.json what the last pull wrote: revisions, content hashes, versions

A case file must be named <id>.json. Commit the folder with your workflow so cases are reviewed in the same pull request as the change they test.

Command Flag Meaning
pull --force Overwrite case files edited locally since the last pull or push
push --cut [note] Cut a dataset version after pushing
push --force Write over cases that changed in the bucket since the last pull
push --prune Delete cases the last pull wrote that you have since removed locally
push --restore Undelete or unarchive cases that are deleted or archived in the bucket but present locally

push writes a new revision only for a case whose content changed. Without --force, it refuses a case that someone edited in the bucket since your last pull, and pull refuses to overwrite a case file you edited locally. Neither removes a case file the last pull did not write. Both refuse a dataset folder whose header.json names another workflow.

Add cases from real runs

A run that went wrong in production is the best case you will get. Capture it instead of retyping it.

In ROC:

  • On a run in Executions, select the Add to dataset button next to Re-run. Pick a Dataset (or type a New dataset name) and confirm the Case id.
  • On the Dataset view, select Add from executions and pick from the workflow’s recent runs.

In code, caseFromRun from @cef-ai/eval builds the same case:

import { addCase, caseFromRun, uniqueCaseId, listCaseIds } from "@cef-ai/eval";
const c = caseFromRun({
input: startPayload, // the run's workflow.start payload
output: completedPayload.output, // workflow.completed output, if the run completed
context: "support-4821",
workflowVersion: "1.3.0",
tags: ["from-production"],
});
c.id = uniqueCaseId(c.id, await listCaseIds(store, "ticket-triage", "triage"));
await addCase(store, "ticket-triage", "triage", c);

What a captured case expects:

The run’s output field is The case expects
A label (60 characters or fewer and 8 words or fewer) That exact value
Free text (longer than that) { "$exists": true }: the field is present
Not an object, or the run did not complete No expectations: state them yourself

An inline graph in the start payload is dropped from the input.

A captured case expects what the workflow did, not what it should have done. Review every captured expectation before you run an experiment, and fix the ones that captured the bug.

Archive and delete

Action Effect
Archive Hides the case from new dataset versions. Its bytes are untouched; versions that pin it still load.
Delete Removes the case from the working set. Its revisions stay (versions may pin them) and its id stays taken.

Both are reversible. cef eval push --restore brings a case back if it is still in your repo folder.

Versions

You rarely cut a version by hand. When an experiment starts without a dataset version, it compares the working set (every live case at its latest revision) with the latest version:

  • unchanged: the experiment runs that version;
  • changed: a new version is cut with the note auto: <n> added, <m> edited, <k> removed, and the experiment runs it.

cef eval push --cut "note" cuts one explicitly. The History table on the Dataset view lists each version with Version, Saved, Note, Changes, Cases, and Experiments.

A version pins each case at one revision by sha256. Revisions are never rewritten, so an experiment from months ago reloads exactly the cases it ran.

Storage

Datasets live in your Agent Service’s bucket, beside the experiments that ran them:

workflows/<wf>/datasets/<dataset>/header.json header (+ baselineId)
workflows/<wf>/datasets/<dataset>/cases/<caseId>/r<n>.json write-once revisions
workflows/<wf>/datasets/<dataset>/cases/<caseId>/deleted.json delete flag
workflows/<wf>/datasets/<dataset>/archived/<caseId>.json archive flag
workflows/<wf>/datasets/<dataset>/versions/v<n>.json frozen versions
workflows/<wf>/experiments/<expId>/… experiments and their runs

<wf> is the workflow’s id (its agent alias). ROC, the CLI, and @cef-ai/eval read and write the same keys.

Anyone who can write a dataset can change what counts as a pass. Treat write access to a dataset like write access to the workflow it tests.