Agents and models
Vision: images and PDFs
Agents that see: send images and PDFs with an agent's input, return screenshots from tools, and keep media checked, limited and out of logs.
An agent can look at images and PDFs: a screenshot of an error, a photo of a receipt, a scanned contract. Media does not travel inside your JSON input. It has its own channel next to the input, its own byte limits, and it is checked before a run starts: its type is read from the file's own bytes, not only from its label.
import { readFile } from 'node:fs/promises';
import { createRuntime, defineAgent, media, z } from 'mayura';
import { openAIResponses } from 'mayura/provider-openai';
const agent = defineAgent({
id: 'receipts', version: '1', instructions: 'Read the receipt and report its total.',
model: openAIResponses({ apiKey, model: modelId, maxCostMicros: 20_000, pricing }),
tools: [],
input: z.object({ question: z.string() }),
output: z.object({ total: z.string(), currency: z.string() }),
media: { accept: ['image/png', 'image/jpeg', 'application/pdf'] },
});
const runtime = createRuntime({
profile: 'ephemeral', permissions: { allow: ['model:openai.responses'] }, limits: { maxCostMicros: 100_000 },
});
const photo = media(await readFile('receipt.jpg'), 'image/jpeg', { name: 'receipt.jpg' });
const outcome = await runtime.submit(agent, { input: { question: 'What is the total?' }, media: [photo] }).result();Media values
| Function | Creates |
|---|---|
media(bytes, mediaType, { name? }) |
Media from bytes (Uint8Array or ArrayBuffer). The bytes are copied, and must really be mediaType: a PNG labelled image/jpeg is refused here. |
mediaUrl(url, mediaType, { name? }) |
Media the model provider fetches itself from an https:// URL. |
mediaFromBase64(text, mediaType, { name? }) |
Media from base64 text, as it arrives in JSON. |
mediaFromArtifact(store, reference, scope) |
Media read from the artifact store, from mayura/artifacts. |
The supported types are image/png, image/jpeg, image/webp, image/gif and application/pdf (MEDIA_TYPES).
Other files, such as SVG or HTML, are never accepted.
What an agent accepts
Nothing is accepted unless the agent declares it with media:
| Field | Meaning |
|---|---|
accept |
The media types allowed. Required. |
maxItems |
How many items one run may carry. Default 4, at most 32. |
maxBytes |
The largest single item, in bytes. Default 10 MB, at most 64 MiB. |
urls |
HTTPS URL prefixes that URL media must start with, each ending in /, such as 'https://cdn.example.com/uploads/'. Without it, only bytes are accepted. |
The runtime limit maxMediaBytes (default 20 MiB) bounds all media in one run, with the input and from tools. It is
counted apart from the JSON limits (maxInputBytes, maxContextBytes), so an image does not crowd out the
conversation.
A refusal happens when you call submit, before the run exists, and says what to change: for example
Agent receipts: media 1 is image/gif, which is not accepted (accepted: image/png, image/jpeg, application/pdf).
URL media is fetched by the model provider, not by Mayura. List only prefixes you trust the provider to read, and prefer bytes for anything private.
Models that can see
A model adapter declares what it can see in capabilities.media. defineAgent refuses an agent whose model cannot
see a type the agent accepts or its tools return, naming the types, so the mistake shows before the first run.
| Adapter | Default | Change it with |
|---|---|---|
openAIResponses |
Every type, and URLs | media: { types, urls }, or media: false for a model that sees nothing |
anthropicMessages |
Every type, and URLs | media: { types, urls }, or media: false |
openAICompatibleChat |
Nothing | media: { types: ['image/png', 'image/jpeg'], urls: true } for a model that can see |
A router sees only what every one of its routes can see. Compatible providers receive images as image_url parts; a
PDF is sent as a file part only if you list application/pdf, only as bytes, and not every provider takes it.
Tools that return images
A tool can return media for the model to look at, such as a screenshot. It declares what it may return with media,
and returns withMedia(output, items):
import { defineTool, media, withMedia, z } from 'mayura';
const screenshot = defineTool({
id: 'browser.screenshot', version: '1', description: 'Take a screenshot of the current page.',
input: z.object({}), output: z.object({ width: z.number(), height: z.number() }),
effects: 'read', capabilities: ['browser'],
media: { accept: ['image/png'], maxBytes: 5_000_000 },
execute: async () => {
const png = await takeScreenshot(); // your browser automation
return withMedia({ width: 1280, height: 800 }, [media(png, 'image/png', { name: 'page.png' })]);
},
});The output is checked against the tool's output schema as usual; the media is checked against the tool's media and
the run's maxMediaBytes. A tool that returns media it did not declare fails with INVALID_OUTPUT. Anthropic models
see the images inside the tool result; OpenAI models see them in a message right after the tool results.
Privacy
- Hooks never see media bytes.
beforeExecutionhasmedia, a summary of each item (type, size, name), andbeforeModelCallhasrequest.media, the same summaries by message index. See Lifecycle hooks. - Events carry only counts:
run.startedandtool.completedhavemediawhen there was some. - Guards check the JSON input and outputs only, so text guards such as
redactPIInever touch image data. - Continuations never store media bytes: they point back to the run's own messages.
Over HTTP
POST /v1/runs takes media next to input: base64 bytes, an allowed URL, or an artifact reference. The client
sends media with submit:
import { createClient } from 'mayura/client';
const client = createClient({ baseUrl: 'https://agents.example.com', token: async () => accessToken });
const file = await fetch('/receipt.jpg').then(response => response.arrayBuffer());
await client.submit('receipts', { question: 'What is the total?' }, {
idempotencyKey: crypto.randomUUID(),
media: [{ mediaType: 'image/jpeg', data: new Uint8Array(file), name: 'receipt.jpg' }],
});The server checks the media against the agent before starting the run and answers 400 INVALID_MEDIA with the exact
reason. See Server and client for the body limit and artifact references.
In workflows
Workflow state holds JSON only. To give a workflow step's agent an image, keep a reference in the step's input and
read the media when the step runs, with agentStep's media option:
import { defineAgent, z } from 'mayura';
import { mediaFromArtifact } from 'mayura/artifacts';
import { agentStep } from 'mayura/workflows/lifecycle';
const receiptAgent = defineAgent({
id: 'receipts', version: '1', instructions: 'Read the receipt and report its total.', model, tools: [],
input: z.object({ receipt: z.record(z.string(), z.unknown()) }), // the stored artifact's reference
output: z.object({ total: z.string() }),
media: { accept: ['image/png', 'image/jpeg', 'application/pdf'] },
});
const readReceipt = agentStep(receiptAgent, {
id: 'receipt.read', permissions: ['model:openai.responses'], limits: { maxCostMicros: 50_000 },
media: async (input, context) => [await mediaFromArtifact(artifacts, input.receipt, context.scope)],
});Testing
scriptedModel can see every type, and a scripted step receives the request with its media. testImage() and
testPdf() from mayura/testing return small valid files. See Testing.