Inference
Rate limits
Direct Inference API request rate, balance, spend-cap, concurrency, and availability limits.
Updated
The direct Inference API at https://api.llmbase.ai/v1 does not use a fixed
requests-per-second or requests-per-minute product quota. Requests are subject
to prepaid balance, optional API-key spend caps, advertised model limits,
availability, and the active-request limit below.
Request rate
No fixed requests-per-second or requests-per-minute quota is enforced for direct Inference API keys. Instead, each account can keep up to 20 prepaid inference requests active at the same time.
That means your completed requests per minute depend on how long your model requests take:
| Average request duration | Approx. completed requests per minute |
|---|---|
| 1 second | 1,200 |
| 5 seconds | 240 |
| 10 seconds | 120 |
| 30 seconds | 40 |
Formula: 20 concurrent requests * 60 / average request duration in seconds.
For example, if your average request takes 5 seconds, 20 concurrent requests
allow about 240 completed requests per minute. If your average request takes 30
seconds, the same limit allows about 40 completed requests per minute.
Active request limit
Each account can have up to 20 active prepaid inference requests reserved at
the same time. If the account already has 20 active prepaid reservations, the
API returns 429 with too_many_concurrent_requests and the message
Too many active inference requests. Retry after current requests finish.
This is an account-level active-request guard, not a model-specific rate limit.
It is separate from Chat and Agent API request limits on llmbase.ai.
Reservations that do not reach model execution expire after 10 minutes. Requests that have started execution remain active until they finish and their usage is settled; this prevents an interrupted client connection from creating duplicate work or an incorrect balance. If a request appears stuck, wait for it to settle before retrying the same workload.
Balance and spend caps
Inference API keys spend prepaid credits. Before a request runs, LLMBase
estimates the maximum cost from the selected model, input size, requested output
tokens, and model output limit. If the account does not have enough available
prepaid balance for that reservation, the API returns 402 with
insufficient_balance.
If an API key has a monthly spend cap, reserved and already-spent cents for that
key are counted against the cap. When a new request would exceed the cap, the API
returns 402 with api_key_spend_cap_exceeded.
Both errors include error.details with the required reservation, the available
amount, and, where possible, a max_tokens value that should fit. See
How billing works for how the reservation is sized
and how to avoid these errors.
Model limits
Model metadata can include context and output limits, but per_request_limits
is null in the public model-list schema. Use each model’s context_length and
max_output_length fields from GET /v1/models?metadata=true to size requests.
When a request exceeds a model context or output limit, reduce the prompt, attachments, or requested output tokens and retry.
Model availability
Temporary failures include a Retry-After header in seconds. Wait at least
that long, then retry with exponential backoff. Common examples:
| Status | error.code | Meaning |
|---|---|---|
429 | model_rate_limited | The selected model is temporarily rate-limited. |
429 | too_many_concurrent_requests | Your account reached the active-request limit above. |
429 | rate_limited | Too many requests in a short time. |
502 | model_response_invalid | The model returned an unusable response; nothing was charged. Retry. |
502 | model_tool_call_incomplete, model_tool_call_text_leak | The model returned an incomplete tool call; nothing was charged. Retry. |
503 | model_temporarily_unavailable | The selected model is temporarily unavailable. |
503 | reservation_unavailable, reservation_dispatch_unavailable | Spend could not be reserved or dispatched right now; nothing was charged. |
503 | admission_unavailable | The API is briefly not accepting new requests. |
504 | model_timeout | The model did not finish in time. Retry with streaming or a smaller request. |
Treat any 429, 502, 503, or 504 with a Retry-After header as
retryable, even if its code is not listed here.
These are separate from the balance and spend-cap limits above, which return
402 and do not succeed on retry.
A streaming request that fails before the first token returns the same HTTP
status and JSON error. If a stream fails after it started, it ends with an
error event whose code is model_stream_error; retry the whole request.
If failures cluster around long-running streams, reduce parallelism. For agent loops, choose the current model from live metadata: start with a lower-cost returned model that supports the features you need, and move to a larger returned model only when context or reasoning requirements demand it.