Test a workflow
Test a workflow on three rungs. Each one catches what the one below cannot.
| Rung | Runs | Catches | When |
|---|---|---|---|
| 1. Local tests | In-process, models mocked | Wrong graph wiring, bad SQL, broken transforms, payloads over budget, bad model output not handled | Every commit |
| 2. Dataset experiments | On the platform, real models | Quality drift: a prompt, model, or step change that makes the Result worse, slower, or costlier | Before switching the live version |
| 3. Live check | The deployed workflow, a real trigger | Deployment, connection, and data problems: the trigger does not fire, a cubby is not written, a widget shows nothing | After every deploy, and daily |
1. Local tests
@cef-ai/testing runs the real workflow runner in-process against an in-memory vault and SQLite cubbies built from your real migrations. You mock only what leaves the workflow: models and other agents.
pnpm add -D @cef-ai/testing vitestThe example is a support-ticket triage workflow: a model classifies the ticket, a transform keeps only known categories, a cubby step records the triage, a publish step tells the support queue, and a Result step ends the run.
ticket-triage/ cef.config.ts cubbies/tickets/001-triage.sql test/triage.test.tsThe workflow
import { fileURLToPath } from "node:url";import { defineWorkflow } from "@cef-ai/agent-sdk/config";
const CATEGORIES = ["billing", "bug", "account", "other"];
export default defineWorkflow({ id: "ticket-triage", version: "1.0.0", models: { classifier: "https://cdn.example.com/models/classifier/model.json" }, cubbies: [{ alias: "tickets", migrations: fileURLToPath(new URL("./cubbies/tickets", import.meta.url)) }], nodes: [ { id: "start", kind: "trigger", label: "Ticket arrives", position: { x: 0, y: 0 }, params: { eventType: "workflow.start" }, }, { id: "classify", kind: "model", label: "Classify the ticket", position: { x: 240, y: 0 }, params: { alias: "classifier", input: { text: "={{ $json.text }}", labels: CATEGORIES }, into: "classification", }, }, { id: "normalize", kind: "transform", label: "Keep only a known category", position: { x: 480, y: 0 }, params: { expr: `({ ticketId: item.ticketId, category: ${JSON.stringify(CATEGORIES)}.includes(item.classification?.category) ? item.classification.category : "other", confidence: typeof item.classification?.confidence === "number" ? item.classification.confidence : 0 })`, }, }, { id: "save", kind: "cubbyExec", label: "Record the triage", position: { x: 720, y: 0 }, params: { alias: "tickets", sql: "INSERT OR IGNORE INTO triage_results (run_id, ticket_id, category, confidence) VALUES (?, ?, ?, ?)", args: ["={{ $runId }}", "={{ $json.ticketId }}", "={{ $json.category }}", "={{ $json.confidence }}"], }, }, { id: "announce", kind: "publish", label: "Tell the support queue", position: { x: 960, y: 0 }, params: { event: "ticket.triaged", payload: JSON.stringify({ ticketId: "", category: "" }), }, }, { id: "result", kind: "output", label: "Result", position: { x: 1200, y: 0 }, params: { result: { category: "={{ $json.category }}", confidence: "={{ $json.confidence }}" }, resultTypes: { category: "string", confidence: "number" }, outcome: "={{ $json.category }}", }, }, ], edges: [ { from: "start", to: "classify" }, { from: "classify", to: "normalize" }, { from: "normalize", to: "save" }, { from: "save", to: "announce" }, { from: "announce", to: "result" }, ],});-- cubbies/tickets/001-triage.sqlCREATE TABLE IF NOT EXISTS triage_results ( run_id TEXT PRIMARY KEY, ticket_id TEXT NOT NULL, category TEXT NOT NULL, confidence REAL NOT NULL);The test
The test registers the workflow runner as an agent with the workflow’s own cubbies and its compiled graph as the deployment param, exactly as the platform does. A run starts with a workflow.start event, the same event an experiment publishes for a case.
import { afterEach, describe, expect, it } from "vitest";import { createModelMock, testPlatform, type ModelMockHandle, type TestPlatform } from "@cef-ai/testing";import { WorkflowRunner, type WorkflowDoc } from "@cef-ai/agent-sdk/workflow";import config from "../cef.config.js";
// The compiled graph, exactly as `cef build` ships it.const graph = JSON.parse((config.params!.graph as { default: string }).default) as WorkflowDoc;const LABELS = ["billing", "bug", "account", "other"];const bytes = (v: unknown) => Buffer.byteLength(JSON.stringify(v ?? null), "utf8");
async function setup(classifier: ModelMockHandle): Promise<TestPlatform> { const p = testPlatform({ agents: { "ticket-triage": { source: WorkflowRunner, cubbies: config.cubbies, params: { graph: config.params!.graph!.default }, }, }, models: { classifier }, }); await p.vault.agents.connect({ agentId: "ticket-triage" }); return p;}
const start = (p: TestPlatform, context: string, payload: Record<string, unknown>) => p.vault.scope("default").publish({ type: "workflow.start", context, payload });
async function events(p: TestPlatform, context: string, type: string) { const page = await p.vault.scope("default").stream(context).events.list<Record<string, unknown>>({ types: [type] }); return page.items.map((e) => e.payload);}
const triageRows = (p: TestPlatform) => p.runInCubby("ticket-triage", "tickets", (cubby) => cubby.query("SELECT * FROM triage_results ORDER BY run_id"));
describe("ticket-triage", () => { let p: TestPlatform | undefined; afterEach(() => p?.dispose());
it("classifies, records one row, publishes a declared payload and ends with the Result", async () => { const classifier = createModelMock(); classifier.expect({ text: "I was charged twice", labels: LABELS }).respond({ category: "billing", confidence: 0.93 }); p = await setup(classifier);
await start(p, "t-1", { ticketId: "T-1", text: "I was charged twice" });
expect(await events(p, "t-1", "workflow.completed")).toEqual([ { runId: "run-t-1", output: { category: "billing", confidence: 0.93 }, outcome: "billing" }, ]); expect(await triageRows(p)).toEqual([ { run_id: "run-t-1", ticket_id: "T-1", category: "billing", confidence: 0.93 }, ]); const [announced] = await events(p, "t-1", "ticket.triaged"); expect(announced).toEqual({ ticketId: "T-1", category: "billing", runId: "run-t-1" }); });
it("downgrades an unknown category to 'other' instead of failing", async () => { const classifier = createModelMock(); classifier.expect({ text: "hello", labels: LABELS }).respond({ category: "spam!!", confidence: "high" }); p = await setup(classifier);
await start(p, "t-2", { ticketId: "T-2", text: "hello" });
const [done] = await events(p, "t-2", "workflow.completed"); expect(done!.output).toEqual({ category: "other", confidence: 0 }); });
it("writes one row when the same start is delivered twice", async () => { const classifier = createModelMock(); classifier.expect({ text: "app crashes", labels: LABELS }).respond({ category: "bug", confidence: 0.8 }); p = await setup(classifier);
await start(p, "t-3", { ticketId: "T-3", text: "app crashes" }); await start(p, "t-3", { ticketId: "T-3", text: "app crashes" });
expect(await triageRows(p)).toHaveLength(1); expect(classifier.calls).toHaveLength(1); });
it("stays inside its budgets", async () => { const text = "x".repeat(20_000); const classifier = createModelMock(); classifier.expect({ text, labels: LABELS }).respond({ category: "bug", confidence: 0.5 }); p = await setup(classifier);
await start(p, "t-4", { ticketId: "T-4", text });
expect(graph.nodes.length).toBeLessThanOrEqual(8); for (const n of graph.nodes) expect(n.position, `${n.id} has a position`).toBeDefined(); expect(classifier.calls).toHaveLength(1); const [done] = await events(p, "t-4", "workflow.completed"); expect(bytes(done!.output)).toBeLessThanOrEqual(1024); const [announced] = await events(p, "t-4", "ticket.triaged"); expect(bytes(announced)).toBeLessThanOrEqual(1024); });});pnpm vitest runWhat each test pins down:
| Test | Guards |
|---|---|
| Happy path | The model is called with the mapped input; the cubby row is exactly right; the publish step sends only its declared fields; the Result and outcome are exact |
| Bad model output | Invalid output is downgraded deterministically, so the run still completes with a valid Result |
| Same start twice | A redelivered start does not run the workflow twice or write two rows |
| Budgets | Step count, positions, model calls per run, Result and payload sizes stay under the numbers you chose |
What else to assert
- Failures. For input the workflow must refuse, assert the run failed and nothing was written: read
workflow.failedfrom the context, and query your cubby for zero rows. The runner’s ownrunscubby holds one row per run withstatusanderror:p.runInCubby("ticket-triage", "runs", (c) => c.query("SELECT status, error FROM runs WHERE run_id = ?", ["run-t-1"])). - Which step wrote what.
p.cubbyOps()lists every cubby statement with itscubby,op,sql,params, and thenodeIdof the step that made it. - Agent steps. Register each agent a step asks as a
mockAgentthat answersworkflow.stepby publishingagent.answeredwith{ agent, text }(or{ agent, error }). See the testing reference. - Publish failures. Set
p.faults.refusePublishto make a publish fail as the vault does for a body over 1 MiB, and assert the run fails visibly. - Fixtures. Keep inputs synthetic: build them with small factory functions in
test/fixtures/, never from real customer data.
Local tests do not prove a real model returns what your mock does. That is the next rung.
2. Dataset experiments
Once the workflow is pushed, run its dataset against the version before it goes live:
cef eval run triage --workflow-version 1.1.0 --repeats 3Start the dataset from the cases your local tests already use (the same start payloads), then grow it with real runs captured from Executions with Add to dataset. Judge every candidate version against a baseline experiment of the live version. See Experiments.
3. Live check
Run this checklist on each deployed workflow after every deploy, and daily on workflows people depend on. It takes a few minutes per workflow.
Prepare (once per workflow)
- Write down a synthetic check input: a start payload that exercises the main path and is safe to run in the vault (no real customer data, no outbound message to a real person).
- Write down what a good run looks like: the expected Result, the cubby row(s) it writes (table, key, columns), the events it publishes, and what the widget shows for it.
- Make sure the check input is idempotent: give it a fixed id so a repeated check overwrites or ignores, not duplicates.
Trigger
- Open the workflow in ROC and confirm the Deployed chip shows the version you expect.
- Start a run the way real traffic does: send the trigger event, call the webhook with an
Idempotency-Key, or wait for the schedule. Use Run in ROC only when the real trigger cannot be fired safely.
Watch the run in Executions
- The run appears at the top of Executions within a few seconds. If it does not, the trigger did not fire: see The run never starts.
- It ends Done. Failed shows the step that failed and why. Waiting names the step it is parked at; that is expected only at a human step.
- Every step you expect ran, in order. Open each model and agent step and check its output looks right.
- Took, Steps, Model calls, Tokens, and Cubby calls are in line with previous runs. A sudden jump is a regression even if the run is Done.
Inspect cubby data
- On a Cubby step of the run, select Open cubby and query the row the run wrote:
SELECT * FROM triage_results WHERE run_id = 'run-<context>'. - Every column you expect is filled; no
NULLwhere a value belongs; timestamps are this run’s. - Exactly one row per run: no duplicates from a retried step.
Verify the widget
- Open the widget where it is pinned and find this run’s record.
- It shows the values in the cubby row, not a placeholder or an empty state.
- Reload the page: the record is still there.
Record
- Note the run id, the version, and pass or fail. On a fail, add the run to the workflow’s dataset with Add to dataset once the fix is known, so an experiment catches it next time.