Skip to content
Mayura

Concepts

Costs and budgets

Cap what a run can spend: costs in micros, per-call ceilings, the run cost limit, tool costs and shared budgets.

Mayura counts money in micros: millionths of a US dollar, as integers. Every model call and tool call declares the most it can cost. Before a call starts, the runtime reserves that amount from the run's budget; after it finishes, it charges the actual cost and releases the rest. A call that does not fit in what is left never starts.

The run budget comes from the runtime's limits.maxCostMicros, and its default is 0. Free calls work without it (the scripted test model costs 0), but a paid model is refused until you set a limit.

A budgeted agent

ts
import { createRuntime, defineAgent, z } from 'mayura';
import { openAIResponses } from 'mayura/provider-openai';

const agent = defineAgent({
  id: 'summarizer',
  version: '1',
  instructions: 'Summarize the text in three sentences.',
  input: z.object({ text: z.string().max(40_000) }),
  output: z.object({ summary: z.string() }),
  tools: [],
  model: openAIResponses({
    apiKey: process.env.OPENAI_API_KEY ?? '',
    model: 'gpt-5-mini',
    // Your model's real prices, in micros per million tokens. $0.25 per million = 250_000.
    pricing: { inputMicrosPerMillionTokens: 250_000, outputMicrosPerMillionTokens: 2_000_000 },
    maxCostMicros: 20_000, // one call may cost at most $0.02
  }),
});

const runtime = createRuntime({
  profile: 'ephemeral',
  permissions: { allow: ['model:openai.responses'] },
  limits: { maxCostMicros: 50_000, maxModelCalls: 2, maxOutputTokens: 1_024 }, // the whole run: at most $0.05
});

const run = runtime.submit(agent, { input: { text: 'A long article to summarize.' } });
const outcome = await run.result();
console.log(outcome.status, runtime.inspect(run).budget);
await runtime.close();

Micros

Amount Micros
$1 1_000_000
1 cent 10_000
$0.001 1_000
1 micro $0.000001

Costs are whole numbers, so there is no rounding drift. Values must be non-negative safe integers.

Model pricing and the per-call ceiling

A provider adapter such as openAIResponses takes the model's prices in micros per million tokens and computes each call's cost from the token counts the provider reports:

cost = (input tokens × input price + output tokens × output price) / 1,000,000, rounded up

The adapter's maxCostMicros is the ceiling for one call. The runtime reserves it before every call, so choose it from the worst case: the largest input you expect plus maxOutputTokens of output. With the prices above, 20,000 input tokens and 4,096 output tokens cost 5,000 + 8,192 = 13,192 micros, so a ceiling of 20,000 leaves room.

Mayura cannot stop a provider from charging more than you expected. If a call's actual cost is above its ceiling, the full cost is still recorded, the run ends as blocked with BUDGET_EXCEEDED, and the run's budget accepts no further calls.

The prices are the ones you configure. Mayura does not look them up, and its totals are an estimate from your prices, not your provider's invoice.

The run cost limit

limits.maxCostMicros caps the total cost of one run, including its child runs. It is per run, not per runtime: every submit starts with a fresh budget of that size.

Before each model call, the runtime checks that the model's maxCostMicros fits in what is left. Before running a batch of tool calls the model asked for, it checks that the costMicros of all of them fits. If not, the run ends as blocked with BUDGET_EXCEEDED, and the call is not made.

So set the run limit to at least the per-call ceiling times the number of model calls you expect, plus your tools' costs. In the example above, two calls of at most 20,000 micros fit in 50,000.

Tool costs

A tool's costMicros (default 0) is the most one call can cost, for tools that spend money themselves: a paid API, an SMS, a search query. It is reserved before the tool runs. If the tool does not say otherwise, a successful call is charged the full amount. To charge less, report the real cost once from inside execute:

ts
import { defineTool, z } from 'mayura';

export const sendSms = defineTool({
  id: 'sms.send',
  version: '1',
  description: 'Send a text message to the signed-in customer.',
  input: z.object({ text: z.string().max(320) }),
  output: z.object({ segments: z.number() }),
  effects: 'write',
  capabilities: ['sms:send'],
  costMicros: 30_000, // at most 3 cents
  execute: async ({ text }, context) => {
    const sent = await sms.send(context.scope.principalId, text, { signal: context.signal });
    context.reportUsage({ knownCostMicros: sent.segments * 7_500, unknownCostMicros: 0 });
    return { segments: sent.segments };
  },
});

Reported cost may not exceed costMicros. Reporting unknownCostMicros above 0 says part of the cost is not settled, and the call ends as outcome_unknown.

When the cost is unknown

Sometimes Mayura cannot know what a call cost: a model call failed without the provider reporting usage, or a tool with effects threw or timed out. Then the call's full ceiling stays reserved. It shows up as reservedMicros and still counts against the budget. A timeout, a cancellation or closing the runtime never turns an unknown cost into zero.

If you write a model adapter, throw ModelInvocationError with the known cost when a call fails after the provider reported usage, so that amount is charged instead of the whole ceiling.

Reading costs

runtime.inspect(run).budget gives the run's totals, including its children, and the run.completed event carries the same numbers:

Field What it is
spentMicros Confirmed cost. A decimal string instead of a number if it ever exceeds the safe integer range.
reservedMicros Held for calls in progress, and for calls whose cost is unknown.
calls Model and tool calls started.

Budgets across runs

The runtime has no budget shared across runs. For a ceiling across many runs or your own calls, Mayura has two building blocks:

  • Budget, from mayura, is an in-memory budget with a cost ceiling and a call ceiling. reserve(maxMicros) throws BUDGET_EXCEEDED if a call does not fit; settle(actual) on the returned reservation records what it cost. fork({ id, maxCostMicros, maxCalls }) makes a child budget that spends from the same funds with a lower ceiling. invokeTool, invokeBatch and runBudgetedTasks from mayura/helpers take a Budget.
  • Durable budgets keep the same kind of ledger in SQLite or PostgreSQL through store.durableBudgets, so it survives restarts. See Storage.
ts
import { Budget } from 'mayura';

const nightly = new Budget(5_000_000, 500); // $5 and 500 calls in total
const perTeam = nightly.fork({ id: 'team-a', maxCostMicros: 1_000_000, maxCalls: 100 });

const reservation = perTeam.reserve(20_000);
reservation.settle(12_400); // the confirmed cost
console.log(nightly.snapshot()); // { spentMicros: 12400, reservedMicros: 0, calls: 1 }

Totals of a budget include its children; don't add a parent's and a child's numbers together.

Good to know

  • Costs are checked before each call. A single call can still overrun its ceiling; Mayura records it and stops the run, but cannot undo the charge.
  • maxModelCalls, maxToolCalls and maxSteps bound a run even when every call is free.
  • Workflows take their own cost limits. See Workflows.