Inference
How billing works
Why the Inference API reserves credit before a request, how the reservation is sized, why a 402 can happen with credit left, and how to avoid it.
Updated
Requests to https://api.llmbase.ai/v1 with an Inference API key are paid from
your prepaid credit balance. This page explains how a request is billed, why
LLMBase reserves credit before it runs, and what to do when a request is
rejected with 402.
1. Two separate budgets
- API credit (prepaid, USD): used by every request to
api.llmbase.aiwith an Inference API key. You top it up on the Balance page, manually or with automatic top-up. - Chat and Agent subscription budget (in %): included in a Chat
subscription and used by Chat and Agents on
llmbase.ai.
The two never mix. API requests with an API key use only your prepaid credit, and Chat subscription budgets do not apply to them. Chat and Agent usage does not reduce your API credit.
2. What happens during a request
- Reserve. Before the model runs, LLMBase holds the maximum amount the request could cost.
- Run. The model processes the request.
- Bill. When it finishes, you are normally charged the actual usage, never more than the reservation.
- Release the rest. The unused part of the reservation is available again immediately.
The reservation is not a charge. It only makes sure your balance can never go negative, even when many requests run at the same time.
3. How the reservation is sized
The reservation is the price of the largest possible request:
- Input: the size of your messages, tools, and other input fields, priced at the model’s input price. The size is counted conservatively, about one token per byte of input, because the exact token count is only known after the request. For normal text this is several times the real token count; the final charge uses the real token count.
- Output limit:
max_tokens(ormax_completion_tokens), priced at the model’s output price. If you do not set it, a default applies:- 16,384 tokens for ordinary requests (or the model’s maximum output, if that is lower);
- for reasoning requests, an output budget that depends on the reasoning effort, by default typically between 5,120 and 28,672 tokens, so the model has room to think and still answer.
n: asking for several choices multiplies the reservation byn.- Images and files:
- An inline image (
data:URL) on a model that accepts images counts as at most 32,768 tokens, and never more than its own size. - A remote image or file URL can point to content of any size, so the request reserves up to the model’s full context window.
- An inline image (
4. Why a 402 can happen while you still have credit
What counts is your available credit: your balance minus the reservations of
your requests that are still running. If the reservation for a new request is
larger than that, LLMBase rejects it with 402 insufficient_balance before any
cost is incurred.
Example. Your balance is $1.50. Ten requests are running and together hold
$1.32, so $0.18 is available. A new request sends about 2,000 bytes of input
(counted as 2,000 tokens for the reservation) to a model priced at $10 per million input tokens and $25 per million output tokens,
without max_tokens:
| Part | Tokens | Reservation |
|---|---|---|
| Input | 2,000 | $0.02 |
| Output (default limit) | 16,384 | $0.41 |
| Total | $0.43 |
$0.43 is more than the $0.18 available, so the request is rejected. The error
suggests max_tokens: 6255, the largest limit that fits right now. If you
expect a short answer, a lower value such as max_tokens: 1200 reserves only
$0.02 + $0.03 = $0.05 and leaves room for other requests.
5. How to avoid a 402
- Set
max_tokensto the answer length you actually expect. This often lowers the reservation by more than 90%. - Limit parallel requests. Each running request holds its own reservation. An account can have up to 20 requests running at the same time; see Rate limits.
- Send images inline instead of as remote URLs on models that accept
images, or set
max_tokenswhen you use remote URLs. On models without image input, inline data is counted by its size and can make the reservation large. - Enable automatic top-up on the Balance page so your balance is refilled before it runs low.
6. What the dashboard shows
- Balance page: your current credit balance, credit activity, and automatic
top-up settings. The balance does not yet subtract the reservations of running
requests; the numbers in a
402response do. - Inference Usage: what your requests actually cost, per period and model. Statistics can lag behind your live balance by a few seconds.
7. The 402 error fields
A rejected reservation returns an OpenAI-compatible error. OpenAI SDKs raise
their usual status error and expose code and type.
{
"error": {
"message": "Insufficient prepaid balance: this request needs a temporary reservation of up to $0.43 ...",
"type": "insufficient_quota",
"param": null,
"code": "insufficient_balance",
"details": {
"required_reservation_usd_cents": 42.96,
"available_usd_cents": 18.0,
"balance_usd_cents": 150.0,
"reserved_by_other_requests_usd_cents": 132.0,
"reserved_max_output_tokens": 16384,
"suggested_max_tokens": 6255,
"auto_topup_triggered": false,
"top_up_url": "https://llmbase.ai/dashboard/balance/",
"docs_url": "https://llmbase.ai/docs/inference/billing"
}
}
}
| Field | Meaning |
|---|---|
required_reservation_usd_cents | Reservation the request needs, in US cents (with fractions) |
available_usd_cents | Balance minus reservations of other running requests |
balance_usd_cents | Your balance, as in GET /v1/balance |
reserved_by_other_requests_usd_cents | Total held by your other running requests (one total for the account) |
reserved_max_output_tokens | Output limit the reservation was priced with; null for embeddings and images |
suggested_max_tokens | A max_tokens value that should fit right now; null if a smaller limit would not help or would be below the reasoning budget (requested effort or the model’s default) |
auto_topup_triggered | Whether an automatic top-up was started by this rejection (currently always false) |
top_up_url, docs_url | Where to add credit and this page |
suggested_max_tokens should fit, but other requests that start before your
retry can still use up the available amount. If it is null, lower the
reasoning effort, shorten the input, or add credit.
When an API key reaches its monthly credit cap, the code is
api_key_spend_cap_exceeded and details contains the key’s values instead of
the account balance: key_spend_cap_usd_cents, key_spent_usd_cents,
key_reserved_usd_cents, available_usd_cents, and api_keys_url. Raise the
cap on the API Keys page or lower max_tokens.
The same format is used by /v1/chat/completions, /v1/completions,
/v1/embeddings, and /v1/images/generations.
8. Interrupted requests
- If a request never starts, its reservation is released after 10 minutes at the latest.
- If the model already responded and your connection drops, you are charged the actual usage. If that cannot be determined, you are charged at most the reserved amount.
- A request that has started can stay active for up to 3 hours until it is settled. Only then does its reservation expire.
9. FAQ
Am I charged the reservation? Normally not: you are charged the actual usage, and the rest of the reservation is released as soon as the request finishes. Only if the actual usage cannot be determined is at most the reserved amount charged (see Interrupted requests).
Why does my balance look high enough?
The balance includes credit that running requests currently hold. Check
available_usd_cents in the 402 details.
Does my Chat subscription cover API requests? No. API requests with an API key use only prepaid credit.
What should I set max_tokens to?
The longest answer you expect. For reasoning models, leave room for the
reasoning effort you request; suggested_max_tokens already does.