@cef-ai/eval
@cef-ai/eval is the library ROC’s Evaluations tab and cef eval are built on. Use it to script datasets and experiments, or to build your own evaluation tooling over the same storage.
pnpm add @cef-ai/eval| Version | 0.1.0 |
| Dependencies | None. Browser-safe: no Node built-ins. |
| Module | ESM |
The library never talks to the network itself. You inject two things: an EvalStore for storage and a WorkflowTarget for starting runs. For concepts, see the Evaluations overview.
Types
Case
interface Case { id: string; // [A-Za-z0-9._-]{1,80} input: Record<string, Json>; // the workflow.start payload expected?: Record<string, Json | Matcher>; // keyed by Result field; dot paths allowed limits?: Limits; tags?: string[]; notes?: string; source?: { context: string; workflowVersion?: string; capturedAt: string };}
interface Limits { maxDurationMs?: number; maxTokens?: number; maxSteps?: number; maxCost?: number }
type Matcher = | { $eq: Json } | { $oneOf: Json[] } | { $contains: Json } | { $regex: string } | { $between: [number, number] } | { $gte: number } | { $lte: number } | { $exists: boolean } | { $approx: number; $tol: number };Matching rules: Results.
Header and VersionManifest
interface Header { dataset: string; workflowId: string; description?: string; limits?: Limits; createdAt: string; baselineId?: string;}
interface VersionManifest { version: number; header: Header; createdAt: string; note?: string; cases: Array<{ id: string; rev: number; sha256: string }>;}Experiment and RunRecord
interface Experiment { id: string; // exp-<base36 time>-<4 chars> dataset: string; datasetVersion: number; workflowId: string; workflowVersion: string; live?: boolean; repeats: number; resultFields?: ResultField[]; attempt?: number; startedAt: string; finishedAt?: string; status: "running" | "done" | "cancelled" | "error"; origin: "roc" | "cli"; createdBy?: string; baselineId?: string; note?: string; error?: string; summary?: Summary;}
interface RunRecord { caseId: string; caseRev: number; repeat: number; attempt?: number; context: string; runId: string; status: "done" | "failed" | "timeout" | "error"; output: Json | null; outcome?: string; error?: string; startedAt: string; endedAt: string | null; metrics: Metrics; scores: Score[]; pass: boolean | null; // null: no expectations, or the harness errored}
interface Metrics { durationMs: number | null; inputTokens: number; outputTokens: number; totalTokens: number; cost: number; steps: number; modelCalls: number; cpuUnits: number; gpuUnits: number;}
interface Score { name: string; // field path, or limit.<name> value: number | string | boolean | null; pass: boolean | null; // null: not judged scorer: string; // "eval/field@1" | "eval/limit@1" source: "eval" | "annotation"; detail?: string;}RunRecord.status: done the workflow completed; failed it announced a failure; timeout it did not end in time; error the harness could not run it.
Summary
interface Summary { cases: number; runs: number; passed: number; failed: number; errored: number; unscored: number; passRate: number | null; // passed / (passed + failed) durationMs: { p50: number | null; p95: number | null; mean: number | null }; tokens: { total: number; meanPerRun: number }; cost: { total: number; meanPerRun: number }; steps: { meanPerRun: number }; limitBreaches: number; fields: Record<string, { passed: number; failed: number; notApplicable: number; rate: number | null }>;}EvalStore
interface EvalStore { list(prefix: string): Promise<string[]>; // every key under prefix, sorted get(key: string): Promise<string | null>; // null when absent; other failures throw put(key: string, body: string): Promise<void>; putIfAbsent(key: string, body: string): Promise<boolean>; // true when this call wrote readonly atomicity?: "conditional-write" | "single-threaded"; listPrefixes?(prefix: string): Promise<string[]>; // optional: immediate child "directories"}| Export | Meaning |
|---|---|
memoryStore(seed?) |
An in-memory store for tests and dry runs, with dump(). |
storeConformance(makeStore) |
A list of { name, run } checks to run against your own implementation (for example, one per it). putIfAbsent must be a real compare-and-set. |
Datasets and cases
Every dataset function takes (store, workflowId, dataset, …).
| Function | Returns | Meaning |
|---|---|---|
createDataset(store, { dataset, workflowId, description?, limits? }) |
Header |
Refuses a name the workflow already has. |
getHeader / requireHeader |
Header | null / Header |
|
updateHeader(store, header) |
Header |
|
setBaseline(store, wf, dataset, expId | null) |
Header |
Set or clear the baseline. |
listWorkflows(store) |
string[] |
Workflows with anything stored. |
listDatasets(store, wf) |
DatasetSummary[] |
{ dataset, cases, latestVersion, experiments, baselineId?, … } |
addCase(store, wf, dataset, c) |
{ rev } |
Refuses an id that exists. |
saveCase(store, wf, dataset, c) |
{ rev, changed, created } |
First revision, next revision, or nothing when identical. |
getCase(store, wf, dataset, id) |
{ case, rev } | null |
Latest revision of a live case. |
getCaseRevision, listCaseRevisions |
One revision; all revision numbers. | |
listCaseIds, listArchivedCaseIds, listDeletedCaseIds |
string[] |
|
archiveCase / unarchiveCase |
Hide from new versions; bytes untouched. | |
deleteCase / undeleteCase |
{ pinnedBy } / |
Tombstone; revisions and the id stay. |
loadWorkingSet(store, wf, dataset) |
{ header, cases } |
Every live case at its latest revision. |
Versions
| Function | Meaning |
|---|---|
ensureDatasetVersion(store, wf, dataset) |
The latest version when the working set matches it; otherwise cuts the next with the note auto: …. Returns { version, created, diff, manifest }. |
cutVersion(store, wf, dataset, { note? }) |
Cut a version explicitly. |
listVersions, latestVersionNumber, getVersion |
Read versions. |
loadVersion(store, wf, dataset, version) |
The cases a version pins, verified by sha256. |
versionHistory(store, wf, dataset) |
Each version with what changed and how many experiments ran it. |
diffVersionEntries(from, to) |
{ added, removed, edited } |
Cases from runs
caseFromRun(run: CapturedRun): Case
interface CapturedRun { input: Record<string, Json>; // workflow.start payload; an inline `graph` is dropped output?: Json | null; // workflow.completed output context: string; workflowVersion?: string; id?: string; // default: derived from the context capturedAt?: string; tags?: string[]; notes?: string;}expected is built from an object output with expectationFrom: labels exactly, free text (isFreeText: over 60 characters or over 8 words) as { $exists: true }. slugCaseId(text) makes a valid id from any text; uniqueCaseId(base, taken) appends -2, -3, … until it is free.
Running experiments
import { runExperiment, resultFieldsFromGraph } from "@cef-ai/eval";
const exp = await runExperiment({ store, workflowId: "ticket-triage", dataset: "triage", workflowVersion: "1.4.0", live: false, // not live: target.pinVersion is required resultFields: resultFieldsFromGraph(graph), repeats: 3, target, onProgress: ({ done, total, run }) => console.log(done, total, run.caseId, run.pass),});RunExperimentOptions
| Option | Default | Meaning |
|---|---|---|
store, workflowId, dataset |
Required. | |
workflowVersion |
Required. The version evaluated (recorded only, when live). |
|
live |
Required. true runs on the live deployment; false needs target.pinVersion. |
|
target |
Required. A WorkflowTarget. |
|
datasetVersion |
ensureDatasetVersion |
Version to run. |
repeats |
1 |
Runs per case. |
concurrency |
4 (DEFAULT_CONCURRENCY) |
Runs in flight. |
timeoutMs |
300000 (DEFAULT_TIMEOUT_MS) |
Per run. |
resultFields |
The version’s declared Result; expectations it rules out score n/a. | |
baselineId, note, origin, createdBy |
Recorded on the experiment. | |
experimentId |
Resume this experiment: runs on file are reused, a new attempt runs the rest. | |
signal |
AbortSignal: stop dispatching; the experiment ends cancelled. |
|
onProgress |
({ done, total, run, reused }) after each run. |
The meta is written first (running), each run record as it finishes, the summary last. Each run publishes into its own context, eval-<expId>-<caseId>-<repeat>-a<attempt>; runOne runs and scores a single case without storage.
WorkflowTarget
interface WorkflowTarget { start(input: Record<string, Json>, context: string): Promise<void>; waitForEnd(context: string, timeoutMs: number): Promise<RunEnd>; metrics(context: string): Promise<Metrics>; pinVersion?(version: string, contextPrefix: string): Promise<() => Promise<void>>;}
interface RunEnd { status: "done" | "failed" | "timeout"; output?: Json; outcome?: string; error?: string; startedAt?: string; endedAt?: string; // server timestamps}start publishes workflow.start with the input into the context; waitForEnd waits for workflow.completed or workflow.failed; metrics reads what the run cost; pinVersion routes contexts starting with the prefix to a version until the returned release function is called.
Reading experiments
| Function | Meaning |
|---|---|
getExperiment(store, wf, expId) |
The experiment, or null. |
listExperiments(store, wf, { dataset? }) |
All experiments of a workflow, or of one dataset. |
loadExperiment(store, wf, expId) |
{ experiment, runs } |
listRunRecords, getRunRecord |
Run records. |
findExperiment(store, expId) |
Find an experiment when the workflow is not known. |
Scoring
| Function | Meaning |
|---|---|
scoreFields(expected, output, resultFields?) |
One eval/field@1 score per expected field. |
scoreLimits(limits, metrics) |
One eval/limit@1 score per limit in force. |
effectiveLimits(header, case) |
The header’s limits, each overridden by the case’s. |
casePass(scores) |
True when every judged field passes; null when none were judged. Limits never count. |
matchValue(expected, actual) |
null on a match, otherwise the reason. |
readField(value, path) |
Read a dot path. |
resultFieldsFromGraph(graph) |
The fields a graph’s Result steps declare (accepts the graph, its JSON text, or a manifest’s params.graph). undefined when none declares a Result. |
checkResultCompatibility(resultFields, cases) |
{ missing, typeChanged, unexpectedNew }: which expectations a version cannot satisfy. |
unsafeRegex(pattern) |
The reason a $regex pattern is refused, or null. MAX_REGEX_LENGTH is 256. |
Summaries and comparison
| Function | Meaning |
|---|---|
summarize(runs) |
A Summary over run records. |
compareExperiments(base, candidate) |
Each side is { experiment?, runs, cases? }. Compares only cases present on both sides at the same revision. |
ExperimentComparison:
| Field | Meaning |
|---|---|
base, candidate |
{ id, workflowVersion?, datasetVersion? } |
comparedOn |
Number of cases compared |
datasetDiff |
{ added, removed, edited } when both sides’ cases are given |
counts |
Runs per change: improved, regressed, unchanged, only-in-base, only-in-candidate |
cases, runs |
Per case and per run: the change, pass rates or statuses, and deltas |
metrics |
{ base, candidate, delta, relative } for pass rate, duration p50 and p95, tokens/run, cost/run, steps/run, limit breaches |
A pass ranks above “not judged”, which ranks above a failure; a case’s change is decided by its mean rank over repeats.
Storage layout helpers
Key builders (headerKey, caseRevKey, versionKey, experimentMetaKey, runKey, …) and validators (isWorkflowId, isDatasetName, isCaseId, isExperimentId and their assert… forms) expose the layout under workflows/<wf>/. canonicalJson and sha256Hex are how versions pin case content. parseCase, parseHeader, parseExperiment, parseRunRecord validate stored JSON and throw EvalContractError.