Structured output
When an agent needs an object rather than prose, constrain the model with response_format, then validate what comes back. Constraining makes a valid result likely; validating makes it certain.
1. Define the shape
import { z } from "zod";
const Verdict = z.object({ verdict: z.enum(["positive", "neutral", "negative"]), score: z.number().int().min(0).max(10), summary: z.string(),});type Verdict = z.infer<typeof Verdict>;2. Constrain generation
Platform LLMs accept response_format and use guided decoding to make the output match it:
response_format |
Constrains output to |
|---|---|
"json" or { type: "json_object" } |
Any valid JSON. |
{ type: "json_schema", schema } |
JSON that matches schema. |
{ type: "json_schema", json_schema: { schema } } |
The same, in the OpenAI request shape. |
Convert the Zod schema with zod-to-json-schema. $refStrategy: "none" inlines definitions so the schema stands alone:
import { zodToJsonSchema } from "zod-to-json-schema";
const responseFormat = { type: "json_schema" as const, schema: zodToJsonSchema(Verdict, { $refStrategy: "none" }),};3. Call, validate, repair once
import type { Context } from "@cef-ai/agent-sdk";import { z } from "zod";import { zodToJsonSchema } from "zod-to-json-schema";
type Message = { role: "system" | "user" | "assistant"; content: string };
export class StructuredOutputError extends Error { constructor(message: string, readonly attempts: number, readonly lastText: string) { super(message); }}
function parseJson(text: string): unknown { return JSON.parse(text.trim().replace(/^```(?:json)?\s*/i, "").replace(/\s*```$/, ""));}
export async function callStructured<T extends z.ZodTypeAny>( ctx: Context, alias: string, schema: T, messages: Message[], maxTokens: number,): Promise<z.infer<T>> { const response_format = { type: "json_schema" as const, schema: zodToJsonSchema(schema, { $refStrategy: "none" }), }; let turn = messages; let lastText = ""; for (let attempt = 1; attempt <= 2; attempt++) { const out = (await ctx.models[alias].infer({ messages: turn, response_format, max_tokens: maxTokens, temperature: 0, })) as { text: string }; lastText = out.text; const parsed = schema.safeParse((() => { try { return parseJson(out.text); } catch { return undefined; } })()); if (parsed.success) return parsed.data; turn = [ ...messages, { role: "assistant", content: out.text }, { role: "user", content: "That did not match the required schema. Reply with only the corrected JSON object." }, ]; } throw new StructuredOutputError("structured output failed after one repair", 2, lastText);}Call it from a handler:
const result = await callStructured(ctx, "llm", Verdict, [ { role: "system", content: "Classify the message. Reply in JSON." }, { role: "user", content: event.payload.text },], 1024);Set max_tokens
The default output budget is 256 tokens. A JSON object cut off at the budget fails to parse with Unexpected end of JSON input, which looks like a schema problem but is a budget problem. Size max_tokens to the largest output your schema allows.
Repair in code before you repair with the model
A repair call costs another inference. Fix mechanical violations in code first, and keep the model repair as the last step:
| Layer | Fix |
|---|---|
| 1. Parse and validate | Zod safeParse. |
| 2. Deterministic repair | Snap to the nearest enum value, clamp numbers into range, trim arrays to their maximum length. Log what you changed. |
| 3. Cross-field checks | Warn on contradictions; do not block. |
| 4. Model repair | One round trip, as above. Never more than one per object. |
When a run extracts several objects, return the ones that succeeded and mark the gaps instead of failing the whole run.
When to use it
Use it when downstream code reads fields from the result. For free text such as summaries or replies, call the model directly and treat the output as a string.