Costs & Limits

Spending in AI is measured not in tokens but in money — and Proxy Agent converts one into the other using the prices you set on the model. So you see not “we spent a million tokens” but “we spent nine dollars”, and you can set a limit in advance.

Two prices for every request

Two independent amounts are recorded for every request. They are not mixed and not averaged — they are simply two different views.

  • Our price — computed from the model prices you set. The spending limits and all of analytics work from it. It is the only amount that affects anything.
  • The vendor’s price — when the vendor itself reports a cost in the reply, it is saved alongside. It serves only for reconciliation: to compare whether your price matches what the vendor charged.

The vendor’s price never limits requests, however large it is. If the vendor reports thousands of dollars while your prices give five cents, the spending limit will count five cents.

How “our” price is computed

The formula is simple: multiply each piece of work by its price per million “portions of text”, add up, divide by a million.

  • The sent text — at the input price.
  • The model’s reply — at the output price. “Reasoning” lands here too: there is no separate price for it, it counts as ordinary output.
  • The reused text — at its own, usually lower, price.
  • Preparing text for reuse — at the temporary-storage write price.

A nuance about reused input: vendors account for it differently in their reports. Some already include it in the total volume of sent text, others show it separately. The product works out by itself whether the report is combined or separate, and never counts the same text twice.

A worked example

A model costs one dollar per million on input and two dollars per million on output. An ordinary short dialogue: 600 “portions of text” sent, 200 came back.

What we countHow muchAt which priceResult
Sent text600$1.00 per million$0.0006
Model’s reply200$2.00 per million$0.0004
Total per request$0.001

Even thousands of such requests cost dollars, not tens of dollars. That is exactly why spending is easier to watch per key: one person or program can quietly eat a noticeable share of the budget without the others noticing.

Special cases

  • A model without prices counts as free. Requests and tokens appear in the log, but the spending stays zero. If you forgot to fill in the prices, the numbers will be understated — worth checking.
  • Blocked requests cost zero. A refusal by spending limit or by the model list adds no spending to any period.
  • Crashed requests also cost zero. When a request dies with an internal error, no spending is booked for it, even if part of the reply had already arrived.
  • A vendor error can cost money. If the vendor answered with an error but its reply carries token data, the spending is counted — the work was done all the same.

The per-key summary

In the “Access Keys” section every key shows how much was spent per day, week, month and in total. The day runs by calendar days, the week from Monday, the month by the calendar. These same amounts are checked before every request, so you cannot quietly exceed a set limit: as soon as spending reaches the limit, the next request is refused.

How to switch spending off

  • Deactivate the key — access closes immediately, no money is spent at all.
  • Set a limit of 0 — the key exists, but not one request passes. A convenient way to “freeze” access without deleting it.
  • Restrict the model list — the key will only reach the cheap models and get refusals for the rest.

Next: what exactly is recorded about every request — in the Request Log section.

← Back to the documentation index