Agents and models
Model providers
Connect an agent to OpenAI, Anthropic or any OpenAI-compatible endpoint, with explicit keys, prices and a per-call cost limit.
A model provider adapter is the object an agent calls to reach a language model. Mayura ships three: OpenAI's
Responses API, Anthropic's Messages API, and a Chat Completions adapter for local servers and the many providers
that speak the OpenAI dialect. Every adapter has the same shape (a ModelAdapter), so agents, tools, guards and
budgets stay the same when you switch provider.
Adapters never look for credentials in the environment, never pick a model or a price for you, never retry, never follow redirects and make no request until an agent calls them. You pass everything explicitly.
import { createRuntime, defineAgent, z } from 'mayura';
import { openAIResponses } from 'mayura/provider-openai';
const Answer = z.object({ answer: z.string() });
const model = openAIResponses({
apiKey: process.env.OPENAI_API_KEY!,
model: process.env.OPENAI_MODEL!,
maxCostMicros: 20_000, // at most $0.02 per model call
// Use your model's current prices, in micros per million tokens ($0.40 = 400_000).
pricing: { inputMicrosPerMillionTokens: 400_000, outputMicrosPerMillionTokens: 1_600_000 },
});
const agent = defineAgent({
id: 'helper', version: '1', instructions: 'Answer the question in one sentence.',
model, tools: [], input: z.object({ question: z.string() }), output: Answer,
});
const runtime = createRuntime({
profile: 'ephemeral',
permissions: { allow: ['model:openai.responses'] },
limits: { maxCostMicros: 100_000 }, // at most $0.10 per run
});
try {
const outcome = await runtime.submit(agent, { input: { question: 'What is a haiku?' } }).result();
console.log(outcome.status === 'succeeded' ? outcome.output.answer : outcome.error);
} finally {
await runtime.close();
}The adapters
| Function | Import from | Adapter id | Streams |
|---|---|---|---|
openAIResponses |
mayura/provider-openai |
openai.responses |
yes |
anthropicMessages |
mayura/provider-anthropic |
anthropic.messages |
yes |
openAICompatibleChat (local) |
mayura/provider-openai |
openai-compatible.chat |
yes |
openAICompatibleChat (remote) |
mayura/provider-openai |
openai-compatible.<id> |
yes |
createModelRouter |
mayura |
the id you give it | yes |
A runtime lets an agent call a model only when its permissions include model:<adapter id>, for example
model:anthropic.messages. A run without that grant ends blocked with PERMISSION_DENIED before any request is
sent. To spread calls over several providers with failover, see Model routing.
openAIResponses sends requests to https://api.openai.com/v1/responses with store: false, and keeps OpenAI's
reasoning items between the model calls of one run. anthropicMessages sends them to
https://api.anthropic.com/v1/messages; it does not turn on prompt caching, extended thinking or provider-hosted
tools. Both destinations are fixed.
import { anthropicMessages } from 'mayura/provider-anthropic';
const claude = anthropicMessages({
apiKey: process.env.ANTHROPIC_API_KEY!,
model: process.env.ANTHROPIC_MODEL!,
maxCostMicros: 30_000,
pricing: { inputMicrosPerMillionTokens: 3_000_000, outputMicrosPerMillionTokens: 15_000_000 },
});Options
All three adapters take these options. openAICompatibleChat adds remote and token, and its apiKey is optional.
openAIResponses and anthropicMessages send to the provider's own API unless you give an endpoint, for example a
gateway (see Through a gateway).
| Option | Required | Meaning |
|---|---|---|
apiKey |
yes | The provider credential. Sent only in the request header. |
model |
yes | The provider's model id (up to 128 characters). |
outputJsonSchema |
no | JSON Schema of the final output. Leave it out: the runtime sends each agent's own (see below). |
maxCostMicros |
yes | The most one model call may cost. Reserved from the run budget before each call. |
pricing |
yes | inputMicrosPerMillionTokens and outputMicrosPerMillionTokens, as non-negative integers. |
timeoutMs |
no | Deadline for one call. Default 30,000. |
maxRequestBytes |
no | Largest request body, not counting images and PDFs (bounded by the run's maxMediaBytes). Default 1 MiB. |
media |
no | What the model can see, { types, urls }. openAIResponses and anthropicMessages see every type and URLs unless told otherwise (false for a model that cannot see); openAICompatibleChat sees nothing unless told. See Vision. |
maxResponseBytes |
no | Largest response body. Default 1 MiB. |
fetch |
no | A fetch-compatible function to send requests through, for tests or a proxy. The destination does not change. |
endpoint |
no | Where requests go instead of the provider's API: an https URL ending in /responses (OpenAI) or /messages (Anthropic), without credentials, query or fragment. |
headers |
no | Extra request headers, such as a gateway's cf-aig-authorization. Treated as credentials. They cannot replace the adapter's own credential and version headers, Content-Type, Host or cookies. With endpoint and a credential header, apiKey may be left out when the gateway holds the provider key. |
Cost is computed from the token usage the provider reports and your prices, rounded up to a whole micro. Anthropic
cache-creation and cache-read tokens are charged at the input price. This is your own accounting, not the provider's
invoice: keep pricing current.
maxCostMicros is a per-call bound. The runtime's run limit (limits.maxCostMicros, default 0) must be at least as
large, or the first call cannot start. If a call reports a cost above its bound, the cost is still recorded and the
run stops with BUDGET_EXCEEDED. The runtime also sends limits.maxOutputTokens (default 4,096) as the provider's
output-token limit, so pick a bound that covers your largest prompt plus that many output tokens. See
Costs and budgets.
JSON Schemas
A provider needs the agent's output, and every tool's input, as JSON Schema. You don't write them: defineAgent and
defineTool generate them from your Zod schemas (any validator that implements Standard JSON Schema works; Zod 4.2
and later does), and the runtime sends the agent's output schema with every call. So one adapter can serve many agents.
Give outputJsonSchema to the adapter, or to defineAgent, only when your validator cannot describe itself.
Adapters ask the provider for strict structured output and strict tool calls, so the schemas must follow the strict rules:
- the root is
type: 'object'; - every object lists all of its properties in
requiredand setsadditionalProperties: false(an object with no properties, such as the input of a tool that takes none, may leaverequiredout).
In Zod terms: make a field the model may leave empty .nullable(), not .optional() or .default(), and avoid
z.record. Transforms are fine; the schema describes what the model writes. defineAgent checks every tool and the
output against these rules with the adapter you give it, and throws INVALID_CONFIG with the field and the fix:
Agent support: The output schema: the object at /properties/tags must set additionalProperties to false; model
providers accept no open objects, records or maps.The agent's output validator still checks the provider's answer after it arrives, so it may be stricter than the
JSON Schema.
Mayura sends each tool under its own id, made safe for provider function names: letters, digits, _ and -, at most
64 characters. So orders.list reaches the model as orders_list, and instructions that name tools by id still
match. Two ids that come out the same (orders.list and orders/list) get _2, _3 and so on. Mayura maps the
names back, and a call to a name that is not one of the agent's tools ends the run. Write descriptions that say what
each tool does and when to use it: the model chooses by name and description.
OpenAI-compatible providers
Many providers and local model servers accept the OpenAI Chat Completions API. openAICompatibleChat talks to them.
Local servers. Without remote, the endpoint must be plain HTTP on localhost, 127.0.0.1 or [::1], with the
exact path /v1/chat/completions. A key is optional.
import { openAICompatibleChat } from 'mayura/provider-openai';
const local = openAICompatibleChat({
endpoint: 'http://127.0.0.1:11434/v1/chat/completions', // for example, Ollama's default port
model: 'my-local-model',
maxCostMicros: 0,
pricing: { inputMicrosPerMillionTokens: 0, outputMicrosPerMillionTokens: 0 },
});Remote providers. Sending prompts to another host is a decision you make explicitly with remote. The endpoint
must be HTTPS and end in /chat/completions, with no credentials in the URL and no query other than api-version.
remote.id (lower-case letters, digits and hyphens) becomes part of the adapter id, so each provider is granted
separately: remote: { id: 'groq' } is granted as model:openai-compatible.groq. A remote provider needs apiKey or
token.
import { openAICompatibleChat } from 'mayura/provider-openai';
const groq = openAICompatibleChat({
endpoint: 'https://api.groq.com/openai/v1/chat/completions',
remote: { id: 'groq' },
apiKey: process.env.GROQ_API_KEY!,
model: process.env.GROQ_MODEL!,
maxCostMicros: 10_000,
pricing: { inputMicrosPerMillionTokens: 100_000, outputMicrosPerMillionTokens: 300_000 },
});Commonly used endpoints (check each provider's current documentation):
| Provider | endpoint |
remote |
|---|---|---|
| Groq | https://api.groq.com/openai/v1/chat/completions |
{ id: 'groq' } |
| Together | https://api.together.xyz/v1/chat/completions |
{ id: 'together' } |
| Fireworks | https://api.fireworks.ai/inference/v1/chat/completions |
{ id: 'fireworks' } |
| Mistral | https://api.mistral.ai/v1/chat/completions |
{ id: 'mistral' } |
| DeepSeek | https://api.deepseek.com/beta/chat/completions (with output: 'json_object', strictTools: true; see below) |
{ id: 'deepseek' } |
| OpenRouter | https://openrouter.ai/api/v1/chat/completions |
{ id: 'openrouter' } |
| xAI | https://api.x.ai/v1/chat/completions |
{ id: 'xai' } |
| Azure OpenAI | https://<resource>.openai.azure.com/openai/deployments/<deployment>/chat/completions?api-version=<version> |
{ id: 'azure', auth: 'api-key' } |
| Gemini | https://generativelanguage.googleapis.com/v1beta/openai/chat/completions |
{ id: 'gemini' } |
| Cloudflare AI Gateway | https://gateway.ai.cloudflare.com/v1/<account>/<gateway>/compat/chat/completions (see below) |
{ id: 'cloudflare' } |
Check your provider with a small budget before you rely on it (pnpm providers:live-check in the Mayura repository does this through the real runtime). Mayura's CI runs against fake servers; DeepSeek and Cloudflare AI Gateway have also passed the live check. The
adapter needs tool calls, structured output and reported token usage (for streaming, stream_options.include_usage).
A provider that answers differently, or leaves out usage, fails the call with a clear reason rather than being guessed
at. Providers that differ from OpenAI in known ways are configured with these options:
| Option | What it is for |
|---|---|
output: 'json_object' |
Providers with JSON mode but no JSON Schema output, such as DeepSeek. The schema goes into the instructions; Mayura still validates the answer. The default is 'json_schema'. |
strictTools: true |
Sends strict: true on every function, for providers with strict tool calls (DeepSeek's /beta endpoint). |
tokenLimitField: 'max_completion_tokens' |
Sends the output-token limit as max_completion_tokens. OpenAI's newer models, directly or as openai/... through a gateway, refuse max_tokens (the default). |
headers |
Extra headers, such as Cloudflare AI Gateway's cf-aig-authorization. Treated as credentials; they cannot replace the key header, Content-Type, Host or cookies. |
body |
Extra request fields a provider defines, such as DeepSeek's thinking. They cannot replace the fields the adapter sets. |
media |
What the model can see, such as { types: ['image/png', 'image/jpeg'], urls: true }. Images are sent as image_url parts; PDFs, if listed, as file parts and only as bytes. See Vision. |
Thinking models. Some providers return the model's reasoning with a tool call and require it back on the next
request (DeepSeek answers 400 without it). The adapter keeps each such turn, reasoning included, in the run's private
state and sends it back, so thinking models can use tools. The reasoning is never released as output or put in
events. anthropicMessages does the same for Claude's signed thinking and redacted_thinking blocks, which Claude
models may return before a tool call or an answer.
DeepSeek. It offers JSON mode but not JSON Schema output, and strict tool calls on its beta endpoint:
import { openAICompatibleChat } from 'mayura/provider-openai';
const deepseek = openAICompatibleChat({
endpoint: 'https://api.deepseek.com/beta/chat/completions',
remote: { id: 'deepseek' },
apiKey: process.env.DEEPSEEK_API_KEY!,
model: 'deepseek-flash',
output: 'json_object',
strictTools: true,
maxCostMicros: 10_000,
pricing, // check DeepSeek's current prices
});Grant it as model:openai-compatible.deepseek. To turn thinking off, add body: { thinking: { type: 'disabled' } }.
Cloudflare AI Gateway fronts many providers with one OpenAI-compatible endpoint. Models are named
provider/model, such as deepseek/deepseek-flash. An authenticated gateway takes its token in cf-aig-authorization;
when the gateway stores the provider's key, no apiKey is needed. Use the dialect of the provider behind it:
import { openAICompatibleChat } from 'mayura/provider-openai';
const gateway = openAICompatibleChat({
endpoint: `https://gateway.ai.cloudflare.com/v1/${accountId}/${gatewayId}/compat/chat/completions`,
remote: { id: 'cloudflare' },
headers: { 'cf-aig-authorization': `Bearer ${process.env.CF_AIG_TOKEN}` },
model: 'deepseek/deepseek-flash',
output: 'json_object', // DeepSeek behind the gateway
strictTools: true,
maxCostMicros: 10_000,
pricing,
});For an OpenAI model behind the gateway (openai/...), leave output at its default, set
tokenLimitField: 'max_completion_tokens', and check the model's own rules: some of OpenAI's reasoning models accept
function tools on Chat Completions only with reasoning turned off, which body: { reasoning_effort: 'none' } does. The
provider's 400 names the setting it refused.
mayura init sets up both: pick DeepSeek or Cloudflare AI Gateway in the wizard.
Through a gateway
openAIResponses and anthropicMessages can send their own protocol through a gateway or proxy, such as Cloudflare AI
Gateway's provider endpoints. Give the full endpoint (not a prefix, whose convention differs between SDKs) and the
gateway's credential in headers. When the gateway stores the provider key, leave apiKey out:
import { anthropicMessages } from 'mayura/provider-anthropic';
import { openAIResponses } from 'mayura/provider-openai';
const base = `https://gateway.ai.cloudflare.com/v1/${accountId}/${gatewayId}`;
const headers = { 'cf-aig-authorization': `Bearer ${process.env.CF_AIG_TOKEN}` };
const gpt = openAIResponses({ endpoint: `${base}/openai/responses`, headers, model: 'gpt-6-luna', maxCostMicros: 10_000, pricing });
const claude = anthropicMessages({ endpoint: `${base}/anthropic/v1/messages`, headers, model: 'claude-sonnet-5', maxCostMicros: 10_000, pricing });The adapters' ids, and so their grants (model:openai.responses, model:anthropic.messages), stay the same. A gateway
is where your prompts and outputs go: choose it as deliberately as the provider. Both adapters pass every
pnpm providers:live-check check through Cloudflare AI Gateway this way.
Azure OpenAI sends the key in the api-key header instead of Authorization: Bearer; set auth: 'api-key'.
Vertex AI uses short-lived Google access tokens. Pass token instead of apiKey: it is called before every
request, so each call gets a fresh token.
const vertex = openAICompatibleChat({
endpoint: vertexChatCompletionsUrl, // Vertex AI's OpenAI-compatible endpoint for your project and region
remote: { id: 'vertex' },
token: () => getAccessToken(), // for example from google-auth-library
model: process.env.VERTEX_MODEL!,
maxCostMicros: 20_000,
pricing,
});The compatible adapter sends the whole conversation on every call. For OpenAI itself, use openAIResponses, which
keeps OpenAI's reasoning items between calls.
Errors
A failed call ends the run with MODEL_FAILED and a message that says why (see Outcomes):
| Reason | For example | Message begins |
|---|---|---|
authentication |
HTTP 401 or 403 | The model provider refused the credentials or access to this model (HTTP 401). |
rate_limited |
HTTP 429, or 402 for quota | The model provider's rate limit or quota was reached (HTTP 429). |
unavailable |
HTTP 5xx or 408, Anthropic's 529, or no connection | The model provider was unavailable. |
timeout |
no answer within timeoutMs |
The model provider did not answer in time. |
rejected |
any other HTTP error, such as 404 for an unknown model | The model provider rejected the request (HTTP 404). |
refused |
the model refused, or stopped at its token limit | The model refused to answer, or stopped before finishing. |
invalid_response |
an answer that is not the required JSON, or no token usage | The model provider returned a response Mayura could not use. |
configuration |
a schema the provider refuses (code INVALID_CONFIG) |
The model adapter could not send this request. |
The messages are Mayura's own. Provider error bodies and your key never appear in outcomes or events. Cancelling a run aborts the request in flight, but cannot prove the provider did no work, so an unconfirmed call keeps its full reservation.
Writing your own adapter
Any object with this shape is a model adapter. Implement one for a provider Mayura does not cover, or to wrap an SDK you already use:
import { MayuraError, type JsonValue, type ModelAdapter, type ModelRequest, type ModelResponse } from 'mayura';
export const myModel: ModelAdapter = {
id: 'my-provider.chat', // granted as model:my-provider.chat
capabilities: { tools: true, structuredOutput: true },
maxCostMicros: 20_000,
async generate(request: ModelRequest): Promise<ModelResponse> {
// request.instructions, request.messages and request.tools describe the call; honour request.signal.
const answer = await callMyProvider(request, { signal: request.signal }) as {
readonly toolCalls: { readonly id: string; readonly toolId: string; readonly input: JsonValue }[];
readonly output: JsonValue;
readonly costMicros: number;
};
if (!answer) throw new MayuraError('MODEL_FAILED', 'The provider returned no answer.');
if (answer.toolCalls.length > 0) {
return { type: 'tool_calls', calls: answer.toolCalls, usage: { costMicros: answer.costMicros } };
}
return { type: 'final', output: answer.output, usage: { costMicros: answer.costMicros } };
},
};generatereturns either tool calls (each with its ownid, the MayuratoolIdand the input) or a final output, with the call's cost in micros. Mayura validates both before using them.- To say why a call failed, throw
new ModelProviderError(reason, { httpStatus, costMicros })with one of the reasons above: the outcome then carries Mayura's message for that reason. PasscostMicroswhen the provider reported usage before failing, so the known cost is charged; otherwise the call keeps its fullmaxCostMicrosreservation. Any other error is reported with a generic message, since an adapter's own text never reaches an outcome. capabilities.media,{ types, urls }, declares the images and PDFs the model can see. Messages then carrymedia(bytes asdata, or aurl);bytesToBase64,mediaDataUrlandencodedMediaBytesfrommayura/core/hosthelp encode them. Leave it out for a model that cannot see. See Vision.- An optional
checkDefinition(definition)method letsdefineAgentcheck the agent's tools and output schema once, so a problem shows when the agent is defined.checkStrictDefinitionfrommayura/core/hostimplements the strict rules above. - An optional
stream(request)method yields{ type: 'output.delta', text }fragments of the final output's JSON text, then exactly one{ type: 'response', response }. See Streaming. continuationon a response is opaque state the runtime hands back on the next call of the same run.
For tests, scriptedModel from mayura/testing is a ready-made adapter; see Testing.
Embeddings
openAIEmbeddings({ apiKey, model, dimensions }) from mayura/provider-openai is an embedding adapter for native
memory, not a chat model. See Memory and context.