Skip to content
Mayura

Agents and models

Streaming

Stream one text field of an agent's answer as the model writes it, with guards on every batch and the final output still checked.

By default Mayura buffers an agent's output: nothing reaches the caller until the whole answer has passed the output schema and every output guard. For a chat, that means the person waits for the full reply. Streaming lets an agent release one text field of its answer, such as reply, while the model is still writing it.

Streamed text is released in batches, and each batch passes your streaming guards before anyone sees it. The final output is still validated and guarded as a whole, and it stays the authoritative result of the run.

ts
import { createRuntime, defineAgent, type Guard, z } from 'mayura';
import { openAIResponses } from 'mayura/provider-openai';

const noCardNumbers: Guard = {
  id: 'no-card-numbers',
  check: value => (/\d{4}[ -]?\d{4}[ -]?\d{4}/u.test(String(value)) ? { decision: 'block' } : { decision: 'allow' }),
};

const agent = defineAgent({
  id: 'support', version: '1', instructions: 'Answer the customer briefly.',
  model: openAIResponses({ apiKey, model: modelId, maxCostMicros: 20_000, pricing }),
  tools: [],
  input: z.object({ message: z.string() }),
  output: z.object({ reply: z.string() }),
  stream: {
    field: ['reply'],        // the string field to stream
    guards: [noCardNumbers], // checked on every batch before it is released
  },
});

const runtime = createRuntime({
  profile: 'ephemeral', permissions: { allow: ['model:openai.responses'] }, limits: { maxCostMicros: 100_000 },
});
const run = runtime.submit(agent, { input: { message: 'Where is my parcel?' } });
for await (const event of run.observe()) {
  if (event.type === 'output.delta') process.stdout.write(String(event.metadata['text']));
}
const outcome = await run.result(); // the validated, guarded answer

The stream policy

Field Meaning
field Path to a string field of the output, 1 to 8 parts: ['reply'], or ['answer', 'text'] for a nested field.
guards Guards run on every batch. Required: pass [] to stream without batch checks, as a deliberate choice.
batch.minChars A batch is released once it has at least this many characters and ends at a space or line break. Default 24.
batch.maxChars A batch never exceeds this many characters. Default 512, at most 4,096.

Smaller batches show text sooner and run the guards more often; larger batches mean fewer checks and more delay. At the end of the answer, whatever is left is released as a last batch.

Batch guards must be ordinary local guards (an object with id and check). Model-backed managed guards are refused here, because they would run on every batch; keep them in the agent's guards.output. See Guardrails.

Guards on batches

Each batch guard receives a string: up to 512 characters of text already released, followed by the new batch. The context lets a guard see a pattern split across two batches, such as a card number cut in half.

  • Allow: the batch is released.
  • Block, or a guard that throws: the batch is withheld, an output.withheld event is emitted, and nothing more is released for that model call. The run goes on: the full answer is still delivered at the end if the final checks pass.
  • Rewrite: return { decision: 'rewrite', value } where value is the same released context followed by the new batch text, for example with an email address masked. The rewritten text is released and the next guard sees it.

The final output stays authoritative

When the model finishes, its complete answer is validated against the agent's output schema and passes the agent's output guards exactly as a buffered answer does. The run's outcome carries that validated output.

Streamed text is provisional, and it is the model's raw text for that field: it has not been through your output schema's transforms. Two consequences:

  • Released text cannot be taken back. If the final checks block the answer, the run fails even though some text was already shown. Put every check that must see each character before anyone does into stream.guards, and keep answers that need a whole-answer decision buffered (no stream).
  • Show the final output when the run completes, and replace the streamed text with it. This matters most when a batch was withheld, or when your output schema rewrites the answer (for example to redact it).

Accounting does not change: the call's cost comes from the complete response. A stream that breaks keeps the call's full reservation unless the provider confirmed its cost.

Consuming the stream

Streamed text arrives as run events of type output.delta, with this metadata:

Key Meaning
step The agent step.
modelCall The model call within the run. A new call starts a new answer.
index The batch's position in that call, from 0.
text The batch text.

output.withheld (with step and modelCall) says a batch guard stopped the stream for that call.

In the same process, read them from run.observe(), as in the example above. The runtime keeps the last 256 events of a run by default (limits.maxEventRetention); an observer that falls further behind receives an events.gap event instead of the missing ones, so observe while the run is in progress.

Over HTTP, the client's headless run store assembles the deltas for you. streamedOutput holds the text of the latest model call, withheld, and complete (false when an event gap dropped text):

ts
import { createHeadlessRunStore } from 'mayura/client/headless';

const store = createHeadlessRunStore({ run: remoteRun });
store.subscribe(() => {
  const state = store.getSnapshot();
  preview.textContent = state.streamedOutput?.text ?? '';
});
await store.observe();

Render streamed text as text (textContent), never as HTML. See Server and client and React. The terminal front ends in mayura/terminal print streamed text as it arrives; see Terminal.

Which models stream

The runtime streams only when the agent has a stream policy and its model adapter has a stream method. openAIResponses, anthropicMessages, openAICompatibleChat and createModelRouter all do. With an adapter that does not, such as scriptedModel from mayura/testing, the agent answers buffered. An agent without a stream policy never streams, whatever its adapter supports.

Only the chosen field is released. Other fields, tool-call arguments, reasoning and raw provider events never are. A model call that ends in tool calls releases no remaining text.

Good to know

  • Mayura's observability records the position and length of each delta, not its text. See Observability.
  • A stream that produces more characters than limits.maxOutputBytes fails the call.
  • With a model router, failover is possible only until the first batch is released.