Inference

How billing works

Why the Inference API reserves credit before a request, how the reservation is sized, why a 402 can happen with credit left, and how to avoid it.

Updated

Requests to https://api.llmbase.ai/v1 with an Inference API key are paid from your prepaid credit balance. This page explains how a request is billed, why LLMBase reserves credit before it runs, and what to do when a request is rejected with 402.

1. Two separate budgets

  • API credit (prepaid, USD): used by every request to api.llmbase.ai with an Inference API key. You top it up on the Balance page, manually or with automatic top-up.
  • Chat and Agent subscription budget (in %): included in a Chat subscription and used by Chat and Agents on llmbase.ai.

The two never mix. API requests with an API key use only your prepaid credit, and Chat subscription budgets do not apply to them. Chat and Agent usage does not reduce your API credit.

2. What happens during a request

  1. Reserve. Before the model runs, LLMBase holds the maximum amount the request could cost.
  2. Run. The model processes the request.
  3. Bill. When it finishes, you are normally charged the actual usage, never more than the reservation.
  4. Release the rest. The unused part of the reservation is available again immediately.

The reservation is not a charge. It only makes sure your balance can never go negative, even when many requests run at the same time.

3. How the reservation is sized

The reservation is the price of the largest possible request:

  • Input: the size of your messages, tools, and other input fields, priced at the model’s input price. The size is counted conservatively, about one token per byte of input, because the exact token count is only known after the request. For normal text this is several times the real token count; the final charge uses the real token count.
  • Output limit: max_tokens (or max_completion_tokens), priced at the model’s output price. If you do not set it, a default applies:
    • 16,384 tokens for ordinary requests (or the model’s maximum output, if that is lower);
    • for reasoning requests, an output budget that depends on the reasoning effort, by default typically between 5,120 and 28,672 tokens, so the model has room to think and still answer.
  • n: asking for several choices multiplies the reservation by n.
  • Images and files:
    • An inline image (data: URL) on a model that accepts images counts as at most 32,768 tokens, and never more than its own size.
    • A remote image or file URL can point to content of any size, so the request reserves up to the model’s full context window.

4. Why a 402 can happen while you still have credit

What counts is your available credit: your balance minus the reservations of your requests that are still running. If the reservation for a new request is larger than that, LLMBase rejects it with 402 insufficient_balance before any cost is incurred.

Example. Your balance is $1.50. Ten requests are running and together hold $1.32, so $0.18 is available. A new request sends about 2,000 bytes of input (counted as 2,000 tokens for the reservation) to a model priced at $10 per million input tokens and $25 per million output tokens, without max_tokens:

PartTokensReservation
Input2,000$0.02
Output (default limit)16,384$0.41
Total$0.43

$0.43 is more than the $0.18 available, so the request is rejected. The error suggests max_tokens: 6255, the largest limit that fits right now. If you expect a short answer, a lower value such as max_tokens: 1200 reserves only $0.02 + $0.03 = $0.05 and leaves room for other requests.

5. How to avoid a 402

  • Set max_tokens to the answer length you actually expect. This often lowers the reservation by more than 90%.
  • Limit parallel requests. Each running request holds its own reservation. An account can have up to 20 requests running at the same time; see Rate limits.
  • Send images inline instead of as remote URLs on models that accept images, or set max_tokens when you use remote URLs. On models without image input, inline data is counted by its size and can make the reservation large.
  • Enable automatic top-up on the Balance page so your balance is refilled before it runs low.

6. What the dashboard shows

  • Balance page: your current credit balance, credit activity, and automatic top-up settings. The balance does not yet subtract the reservations of running requests; the numbers in a 402 response do.
  • Inference Usage: what your requests actually cost, per period and model. Statistics can lag behind your live balance by a few seconds.

7. The 402 error fields

A rejected reservation returns an OpenAI-compatible error. OpenAI SDKs raise their usual status error and expose code and type.

{
  "error": {
    "message": "Insufficient prepaid balance: this request needs a temporary reservation of up to $0.43 ...",
    "type": "insufficient_quota",
    "param": null,
    "code": "insufficient_balance",
    "details": {
      "required_reservation_usd_cents": 42.96,
      "available_usd_cents": 18.0,
      "balance_usd_cents": 150.0,
      "reserved_by_other_requests_usd_cents": 132.0,
      "reserved_max_output_tokens": 16384,
      "suggested_max_tokens": 6255,
      "auto_topup_triggered": false,
      "top_up_url": "https://llmbase.ai/dashboard/balance/",
      "docs_url": "https://llmbase.ai/docs/inference/billing"
    }
  }
}
FieldMeaning
required_reservation_usd_centsReservation the request needs, in US cents (with fractions)
available_usd_centsBalance minus reservations of other running requests
balance_usd_centsYour balance, as in GET /v1/balance
reserved_by_other_requests_usd_centsTotal held by your other running requests (one total for the account)
reserved_max_output_tokensOutput limit the reservation was priced with; null for embeddings and images
suggested_max_tokensA max_tokens value that should fit right now; null if a smaller limit would not help or would be below the reasoning budget (requested effort or the model’s default)
auto_topup_triggeredWhether an automatic top-up was started by this rejection (currently always false)
top_up_url, docs_urlWhere to add credit and this page

suggested_max_tokens should fit, but other requests that start before your retry can still use up the available amount. If it is null, lower the reasoning effort, shorten the input, or add credit.

When an API key reaches its monthly credit cap, the code is api_key_spend_cap_exceeded and details contains the key’s values instead of the account balance: key_spend_cap_usd_cents, key_spent_usd_cents, key_reserved_usd_cents, available_usd_cents, and api_keys_url. Raise the cap on the API Keys page or lower max_tokens.

The same format is used by /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/images/generations.

8. Interrupted requests

  • If a request never starts, its reservation is released after 10 minutes at the latest.
  • If the model already responded and your connection drops, you are charged the actual usage. If that cannot be determined, you are charged at most the reserved amount.
  • A request that has started can stay active for up to 3 hours until it is settled. Only then does its reservation expire.

9. FAQ

Am I charged the reservation? Normally not: you are charged the actual usage, and the rest of the reservation is released as soon as the request finishes. Only if the actual usage cannot be determined is at most the reserved amount charged (see Interrupted requests).

Why does my balance look high enough? The balance includes credit that running requests currently hold. Check available_usd_cents in the 402 details.

Does my Chat subscription cover API requests? No. API requests with an API key use only prepaid credit.

What should I set max_tokens to? The longest answer you expect. For reasoning models, leave room for the reasoning effort you request; suggested_max_tokens already does.