Skip to content

Test a workflow

Test a workflow on three rungs. Each one catches what the one below cannot.

Rung Runs Catches When
1. Local tests In-process, models mocked Wrong graph wiring, bad SQL, broken transforms, payloads over budget, bad model output not handled Every commit
2. Dataset experiments On the platform, real models Quality drift: a prompt, model, or step change that makes the Result worse, slower, or costlier Before switching the live version
3. Live check The deployed workflow, a real trigger Deployment, connection, and data problems: the trigger does not fire, a cubby is not written, a widget shows nothing After every deploy, and daily

1. Local tests

@cef-ai/testing runs the real workflow runner in-process against an in-memory vault and SQLite cubbies built from your real migrations. You mock only what leaves the workflow: models and other agents.

Terminal window
pnpm add -D @cef-ai/testing vitest

The example is a support-ticket triage workflow: a model classifies the ticket, a transform keeps only known categories, a cubby step records the triage, a publish step tells the support queue, and a Result step ends the run.

ticket-triage/
cef.config.ts
cubbies/tickets/001-triage.sql
test/triage.test.ts

The workflow

cef.config.ts
import { fileURLToPath } from "node:url";
import { defineWorkflow } from "@cef-ai/agent-sdk/config";
const CATEGORIES = ["billing", "bug", "account", "other"];
export default defineWorkflow({
id: "ticket-triage",
version: "1.0.0",
models: { classifier: "https://cdn.example.com/models/classifier/model.json" },
cubbies: [{ alias: "tickets", migrations: fileURLToPath(new URL("./cubbies/tickets", import.meta.url)) }],
nodes: [
{
id: "start",
kind: "trigger",
label: "Ticket arrives",
position: { x: 0, y: 0 },
params: { eventType: "workflow.start" },
},
{
id: "classify",
kind: "model",
label: "Classify the ticket",
position: { x: 240, y: 0 },
params: {
alias: "classifier",
input: { text: "={{ $json.text }}", labels: CATEGORIES },
into: "classification",
},
},
{
id: "normalize",
kind: "transform",
label: "Keep only a known category",
position: { x: 480, y: 0 },
params: {
expr: `({
ticketId: item.ticketId,
category: ${JSON.stringify(CATEGORIES)}.includes(item.classification?.category)
? item.classification.category : "other",
confidence: typeof item.classification?.confidence === "number"
? item.classification.confidence : 0
})`,
},
},
{
id: "save",
kind: "cubbyExec",
label: "Record the triage",
position: { x: 720, y: 0 },
params: {
alias: "tickets",
sql: "INSERT OR IGNORE INTO triage_results (run_id, ticket_id, category, confidence) VALUES (?, ?, ?, ?)",
args: ["={{ $runId }}", "={{ $json.ticketId }}", "={{ $json.category }}", "={{ $json.confidence }}"],
},
},
{
id: "announce",
kind: "publish",
label: "Tell the support queue",
position: { x: 960, y: 0 },
params: {
event: "ticket.triaged",
payload: JSON.stringify({ ticketId: "", category: "" }),
},
},
{
id: "result",
kind: "output",
label: "Result",
position: { x: 1200, y: 0 },
params: {
result: { category: "={{ $json.category }}", confidence: "={{ $json.confidence }}" },
resultTypes: { category: "string", confidence: "number" },
outcome: "={{ $json.category }}",
},
},
],
edges: [
{ from: "start", to: "classify" },
{ from: "classify", to: "normalize" },
{ from: "normalize", to: "save" },
{ from: "save", to: "announce" },
{ from: "announce", to: "result" },
],
});
-- cubbies/tickets/001-triage.sql
CREATE TABLE IF NOT EXISTS triage_results (
run_id TEXT PRIMARY KEY,
ticket_id TEXT NOT NULL,
category TEXT NOT NULL,
confidence REAL NOT NULL
);

The test

The test registers the workflow runner as an agent with the workflow’s own cubbies and its compiled graph as the deployment param, exactly as the platform does. A run starts with a workflow.start event, the same event an experiment publishes for a case.

test/triage.test.ts
import { afterEach, describe, expect, it } from "vitest";
import { createModelMock, testPlatform, type ModelMockHandle, type TestPlatform } from "@cef-ai/testing";
import { WorkflowRunner, type WorkflowDoc } from "@cef-ai/agent-sdk/workflow";
import config from "../cef.config.js";
// The compiled graph, exactly as `cef build` ships it.
const graph = JSON.parse((config.params!.graph as { default: string }).default) as WorkflowDoc;
const LABELS = ["billing", "bug", "account", "other"];
const bytes = (v: unknown) => Buffer.byteLength(JSON.stringify(v ?? null), "utf8");
async function setup(classifier: ModelMockHandle): Promise<TestPlatform> {
const p = testPlatform({
agents: {
"ticket-triage": {
source: WorkflowRunner,
cubbies: config.cubbies,
params: { graph: config.params!.graph!.default },
},
},
models: { classifier },
});
await p.vault.agents.connect({ agentId: "ticket-triage" });
return p;
}
const start = (p: TestPlatform, context: string, payload: Record<string, unknown>) =>
p.vault.scope("default").publish({ type: "workflow.start", context, payload });
async function events(p: TestPlatform, context: string, type: string) {
const page = await p.vault.scope("default").stream(context).events.list<Record<string, unknown>>({ types: [type] });
return page.items.map((e) => e.payload);
}
const triageRows = (p: TestPlatform) =>
p.runInCubby("ticket-triage", "tickets", (cubby) => cubby.query("SELECT * FROM triage_results ORDER BY run_id"));
describe("ticket-triage", () => {
let p: TestPlatform | undefined;
afterEach(() => p?.dispose());
it("classifies, records one row, publishes a declared payload and ends with the Result", async () => {
const classifier = createModelMock();
classifier.expect({ text: "I was charged twice", labels: LABELS }).respond({ category: "billing", confidence: 0.93 });
p = await setup(classifier);
await start(p, "t-1", { ticketId: "T-1", text: "I was charged twice" });
expect(await events(p, "t-1", "workflow.completed")).toEqual([
{ runId: "run-t-1", output: { category: "billing", confidence: 0.93 }, outcome: "billing" },
]);
expect(await triageRows(p)).toEqual([
{ run_id: "run-t-1", ticket_id: "T-1", category: "billing", confidence: 0.93 },
]);
const [announced] = await events(p, "t-1", "ticket.triaged");
expect(announced).toEqual({ ticketId: "T-1", category: "billing", runId: "run-t-1" });
});
it("downgrades an unknown category to 'other' instead of failing", async () => {
const classifier = createModelMock();
classifier.expect({ text: "hello", labels: LABELS }).respond({ category: "spam!!", confidence: "high" });
p = await setup(classifier);
await start(p, "t-2", { ticketId: "T-2", text: "hello" });
const [done] = await events(p, "t-2", "workflow.completed");
expect(done!.output).toEqual({ category: "other", confidence: 0 });
});
it("writes one row when the same start is delivered twice", async () => {
const classifier = createModelMock();
classifier.expect({ text: "app crashes", labels: LABELS }).respond({ category: "bug", confidence: 0.8 });
p = await setup(classifier);
await start(p, "t-3", { ticketId: "T-3", text: "app crashes" });
await start(p, "t-3", { ticketId: "T-3", text: "app crashes" });
expect(await triageRows(p)).toHaveLength(1);
expect(classifier.calls).toHaveLength(1);
});
it("stays inside its budgets", async () => {
const text = "x".repeat(20_000);
const classifier = createModelMock();
classifier.expect({ text, labels: LABELS }).respond({ category: "bug", confidence: 0.5 });
p = await setup(classifier);
await start(p, "t-4", { ticketId: "T-4", text });
expect(graph.nodes.length).toBeLessThanOrEqual(8);
for (const n of graph.nodes) expect(n.position, `${n.id} has a position`).toBeDefined();
expect(classifier.calls).toHaveLength(1);
const [done] = await events(p, "t-4", "workflow.completed");
expect(bytes(done!.output)).toBeLessThanOrEqual(1024);
const [announced] = await events(p, "t-4", "ticket.triaged");
expect(bytes(announced)).toBeLessThanOrEqual(1024);
});
});
Terminal window
pnpm vitest run

What each test pins down:

Test Guards
Happy path The model is called with the mapped input; the cubby row is exactly right; the publish step sends only its declared fields; the Result and outcome are exact
Bad model output Invalid output is downgraded deterministically, so the run still completes with a valid Result
Same start twice A redelivered start does not run the workflow twice or write two rows
Budgets Step count, positions, model calls per run, Result and payload sizes stay under the numbers you chose

What else to assert

  • Failures. For input the workflow must refuse, assert the run failed and nothing was written: read workflow.failed from the context, and query your cubby for zero rows. The runner’s own runs cubby holds one row per run with status and error: p.runInCubby("ticket-triage", "runs", (c) => c.query("SELECT status, error FROM runs WHERE run_id = ?", ["run-t-1"])).
  • Which step wrote what. p.cubbyOps() lists every cubby statement with its cubby, op, sql, params, and the nodeId of the step that made it.
  • Agent steps. Register each agent a step asks as a mockAgent that answers workflow.step by publishing agent.answered with { agent, text } (or { agent, error }). See the testing reference.
  • Publish failures. Set p.faults.refusePublish to make a publish fail as the vault does for a body over 1 MiB, and assert the run fails visibly.
  • Fixtures. Keep inputs synthetic: build them with small factory functions in test/fixtures/, never from real customer data.

Local tests do not prove a real model returns what your mock does. That is the next rung.

2. Dataset experiments

Once the workflow is pushed, run its dataset against the version before it goes live:

Terminal window
cef eval run triage --workflow-version 1.1.0 --repeats 3

Start the dataset from the cases your local tests already use (the same start payloads), then grow it with real runs captured from Executions with Add to dataset. Judge every candidate version against a baseline experiment of the live version. See Experiments.

3. Live check

Run this checklist on each deployed workflow after every deploy, and daily on workflows people depend on. It takes a few minutes per workflow.

Prepare (once per workflow)

  • Write down a synthetic check input: a start payload that exercises the main path and is safe to run in the vault (no real customer data, no outbound message to a real person).
  • Write down what a good run looks like: the expected Result, the cubby row(s) it writes (table, key, columns), the events it publishes, and what the widget shows for it.
  • Make sure the check input is idempotent: give it a fixed id so a repeated check overwrites or ignores, not duplicates.

Trigger

  • Open the workflow in ROC and confirm the Deployed chip shows the version you expect.
  • Start a run the way real traffic does: send the trigger event, call the webhook with an Idempotency-Key, or wait for the schedule. Use Run in ROC only when the real trigger cannot be fired safely.

Watch the run in Executions

  • The run appears at the top of Executions within a few seconds. If it does not, the trigger did not fire: see The run never starts.
  • It ends Done. Failed shows the step that failed and why. Waiting names the step it is parked at; that is expected only at a human step.
  • Every step you expect ran, in order. Open each model and agent step and check its output looks right.
  • Took, Steps, Model calls, Tokens, and Cubby calls are in line with previous runs. A sudden jump is a regression even if the run is Done.

Inspect cubby data

  • On a Cubby step of the run, select Open cubby and query the row the run wrote: SELECT * FROM triage_results WHERE run_id = 'run-<context>'.
  • Every column you expect is filled; no NULL where a value belongs; timestamps are this run’s.
  • Exactly one row per run: no duplicates from a retried step.

Verify the widget

  • Open the widget where it is pinned and find this run’s record.
  • It shows the values in the cubby row, not a placeholder or an empty state.
  • Reload the page: the record is still there.

Record

  • Note the run id, the version, and pass or fail. On a fail, add the run to the workflow’s dataset with Add to dataset once the fix is known, so an experiment catches it next time.