Skip to content

Models

Building blocks are what your agents use besides their own logic: models, cubbies, and widgets. Each belongs to your Agent Service, and each page here shows how a workflow and a code agent use it.

Agents and workflows call models without holding an API key, picking a provider, or knowing where a GPU lives. You bind an alias to a model from the platform’s catalogue, call the alias, and the platform routes the request to a server that is serving that model, meters the call, and returns the output.

Where models come from

The platform serves a curated catalogue. Each model is described by a model.json in content-addressed storage, at:

<cdn>/<bucket>/models/<name>/<version>/model.json

The model.json names the model, its alias, its tasks, its input and output schema, and whether it streams. The catalogue spans text generation, speech-to-text, embeddings, vision, and other tasks. What a model accepts and returns is defined by its own schema, not by a fixed request shape.

Open Models in ROC’s side navigation to browse it. Each card shows the model’s name, version, task, alias (as ctx.models.<alias>), price, and whether it is available; search by name, alias, task, or version, and filter by availability (Active or Pending). Open a model to see its input and output schema and try it.

The models in your own Agent Service’s bucket are also listed by the API:

GET /api/v1/agent-services/:asPubKey/models

Declare an alias

A workflow and a code agent declare models the same way: a models map in cef.config.ts, from alias to the model’s model.json URL.

import { defineAgent } from "@cef-ai/agent-sdk/config";
export default defineAgent({
id: "summarizer",
version: "1.0.0",
entry: "./src/agent.ts",
models: {
llm: "https://cdn.ddc-dragon.com/<bucket>/models/<name>/<version>/model.json",
},
});

defineWorkflow takes the same models field.

Rule Detail
URL The path must be /<bucket>/models/<name>/<version>/model.json, with an integer bucket. cef build reads the bucket, name, and version from it and refuses any other shape.
Key Must equal the alias written inside that model.json, which is often not the name in its URL. A key that does not match fails at the first call with model alias not declared: '<alias>'.
Use A model step or a ctx.models.<alias> reference to an alias that is not in models fails cef build.

The alias is the stable handle; the binding lives in config. Swapping a model is a config change and a redeploy, not a code change.

In a workflow

A model step names an alias, resolves its request from the carried item, calls the model once, and puts the answer on the item under into:

{
id: "classify",
kind: "model",
params: {
alias: "llm",
input: { prompt: "=Classify this ticket:\n{{ $json.text }}", max_tokens: 64 },
into: "classification",
},
}

Its parameters, structured output, per-entry calls, and when to use an Agent step instead are on Steps: Model.

In a code agent

A code agent calls a declared model with ctx.models[alias].infer(input).

Type it. Run cef typegen after you declare or change models. It reads each declared model.json and writes .cef/generated.d.ts, which types ctx.models.<alias> from the model’s input and output schemas, and it writes cef.lock.json.

Call it.

@OnEvent("document.added")
async onDocument(event: Event<{ text: string }>, ctx: Context) {
const out = (await ctx.models.llm.infer({
messages: [
{ role: "system", content: "Summarize in two sentences." },
{ role: "user", content: event.payload.text },
],
temperature: 0,
max_tokens: 400,
})) as { text: string };
await ctx.vault.publish("document.summarized", { summary: out.text });
}
Method Behavior
infer(input) Sends input to the model and resolves with the model’s output. The input shape is the model’s own, as declared in its model.json.
stream(input) Returns an async iterable. In the agent sandbox it yields the complete output once. It does not stream tokens.

To pass a vault object to a model that fetches it, mint a short-lived URL with ctx.vault.objects.presignedUrl(path) and send the URL.

Choose the model per deployment. Declare a modelAlias param so a deployment can switch models without a rebuild:

models: {
small: "https://cdn.ddc-dragon.com/<bucket>/models/<small>/<version>/model.json",
large: "https://cdn.ddc-dragon.com/<bucket>/models/<large>/<version>/model.json",
},
params: {
model: { type: "modelAlias", default: "small", enum: ["small", "large"] },
},
const alias = String(ctx.params.model);
const out = await ctx.models[alias].infer({ prompt: event.payload.text, max_tokens: 400 });

enum must list declared aliases. A deployment that sets a value outside enum falls back to the default. See Push and deploy.

Test without a cluster. Stub the alias in the test harness:

import { testAgent, createModelMock } from "@cef-ai/testing";
const h = testAgent(Summarizer, {
models: { llm: { infer: async () => "A short summary.", stream: async function* () {} } },
});

createModelMock() returns a handle with .expect(input).respond(output) and a .calls array for assertions. See the testing reference.

Write the input

The input is the model’s own schema; read it on the model’s page in ROC, or rely on the types cef typegen generates. An ASR model, for example, takes an audio URL.

The platform’s language models take this input and return { text }, plus tool_calls when the model calls a tool and usage when the model reports it:

Field Default Meaning
messages — Chat messages { role, content }. Either messages or prompt is required.
prompt — A single user prompt.
max_tokens 256 Output budget. Set it explicitly: long output, JSON in particular, is cut off at the budget.
temperature 0.7 Sampling temperature.
top_p, top_k, stop — Standard sampling controls.
frequency_penalty, presence_penalty, repetition_penalty — Repetition controls.
response_format — Constrain output: "json", { type: "json_object" }, or { type: "json_schema", schema }. See Structured output.
tools — Tool declarations, rendered into the prompt by the model’s chat template.
image — Image URL, for vision models.

Metering

Every inference call is metered against the run that made it. Each call records its duration and, when the model reports them, input tokens, output tokens, and GPU units on the Task. GPU units count toward the connection’s Compute limit, which the vault owner can set; once it is reached, new work stops. See Spend limits.

Failures

Inference crosses the network. A model can be pending, busy, or slow.

  • In a workflow, a model call that fails in transport (a 5xx, a timeout, a dropped connection) is retried up to 3 attempts in total, 0.4 s and then 1.6 s apart. A 4xx fails the step at once. Treat a failed model step as a normal failure path.
  • In a code agent, a failed call rejects infer. Catch it and degrade gracefully, and write a model’s output to a cubby as soon as you have it, so a retried Task does not pay for the same call twice.

In tests, createModelMock() from @cef-ai/testing scripts a model’s answers for both. See Test a workflow and Test and debug.

Limits

Limit Value
Model call attempts in a workflow 3, transport failures only, 400 ms then 1600 ms apart