Skip to content

Structured output

When an agent needs an object rather than prose, constrain the model with response_format, then validate what comes back. Constraining makes a valid result likely; validating makes it certain.

1. Define the shape

import { z } from "zod";
const Verdict = z.object({
verdict: z.enum(["positive", "neutral", "negative"]),
score: z.number().int().min(0).max(10),
summary: z.string(),
});
type Verdict = z.infer<typeof Verdict>;

2. Constrain generation

Platform LLMs accept response_format and use guided decoding to make the output match it:

response_format Constrains output to
"json" or { type: "json_object" } Any valid JSON.
{ type: "json_schema", schema } JSON that matches schema.
{ type: "json_schema", json_schema: { schema } } The same, in the OpenAI request shape.

Convert the Zod schema with zod-to-json-schema. $refStrategy: "none" inlines definitions so the schema stands alone:

import { zodToJsonSchema } from "zod-to-json-schema";
const responseFormat = {
type: "json_schema" as const,
schema: zodToJsonSchema(Verdict, { $refStrategy: "none" }),
};

3. Call, validate, repair once

import type { Context } from "@cef-ai/agent-sdk";
import { z } from "zod";
import { zodToJsonSchema } from "zod-to-json-schema";
type Message = { role: "system" | "user" | "assistant"; content: string };
export class StructuredOutputError extends Error {
constructor(message: string, readonly attempts: number, readonly lastText: string) {
super(message);
}
}
function parseJson(text: string): unknown {
return JSON.parse(text.trim().replace(/^```(?:json)?\s*/i, "").replace(/\s*```$/, ""));
}
export async function callStructured<T extends z.ZodTypeAny>(
ctx: Context,
alias: string,
schema: T,
messages: Message[],
maxTokens: number,
): Promise<z.infer<T>> {
const response_format = {
type: "json_schema" as const,
schema: zodToJsonSchema(schema, { $refStrategy: "none" }),
};
let turn = messages;
let lastText = "";
for (let attempt = 1; attempt <= 2; attempt++) {
const out = (await ctx.models[alias].infer({
messages: turn,
response_format,
max_tokens: maxTokens,
temperature: 0,
})) as { text: string };
lastText = out.text;
const parsed = schema.safeParse((() => { try { return parseJson(out.text); } catch { return undefined; } })());
if (parsed.success) return parsed.data;
turn = [
...messages,
{ role: "assistant", content: out.text },
{ role: "user", content: "That did not match the required schema. Reply with only the corrected JSON object." },
];
}
throw new StructuredOutputError("structured output failed after one repair", 2, lastText);
}

Call it from a handler:

const result = await callStructured(ctx, "llm", Verdict, [
{ role: "system", content: "Classify the message. Reply in JSON." },
{ role: "user", content: event.payload.text },
], 1024);

Set max_tokens

The default output budget is 256 tokens. A JSON object cut off at the budget fails to parse with Unexpected end of JSON input, which looks like a schema problem but is a budget problem. Size max_tokens to the largest output your schema allows.

Repair in code before you repair with the model

A repair call costs another inference. Fix mechanical violations in code first, and keep the model repair as the last step:

Layer Fix
1. Parse and validate Zod safeParse.
2. Deterministic repair Snap to the nearest enum value, clamp numbers into range, trim arrays to their maximum length. Log what you changed.
3. Cross-field checks Warn on contradictions; do not block.
4. Model repair One round trip, as above. Never more than one per object.

When a run extracts several objects, return the ones that succeeded and mark the gaps instead of failing the whole run.

When to use it

Use it when downstream code reads fields from the result. For free text such as summaries or replies, call the model directly and treat the output as a string.