
LLM API price tables show a rate, not the cost of a completed task. Compare input, cached input, output, tool charges, model behavior, and billing terms with the same workload.
This guide compares public products that developers can buy directly from each company. Prices below are representative public list prices checked on 31 August 2026, shown in USD per one million text tokens unless noted. Vendors can change prices and model availability. Open the linked primary source before making a purchase or production decision.
Representative public prices
| Public API and model | Input | Cached input or cache read | Output | Primary source checked 31 Aug 2026 |
|---|---|---|---|---|
| LLMBase, DeepSeek V4 Flash | $0.20 | $0.07 | $0.30 | LLMBase live model catalog |
| OpenAI, GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | OpenAI model documentation |
| Anthropic, Sonnet 5 | $2.00 | $0.20 cache read | $10.00 | Anthropic API pricing |
| Google, Gemini 3.5 Flash | $1.50 | $0.15 plus cache storage | $9.00 | Gemini API pricing |
| Mistral, Mistral Large | $0.50 | Up to 90% lower for cached input | $1.50 | Mistral pricing |
These models do not have equal quality, speed, context, or tool support. The table illustrates pricing structures. It does not rank model capability.
Caching also means different things across APIs. Check whether a vendor charges for cache writes, reads, storage duration, or automatic reuse. A low cache-read rate has little value when your prompts rarely share a stable prefix.
Calculate the cost of your workload
Use measured token counts from a representative request. A basic monthly estimate is:
uncached input tokens × input rate
+ cached tokens × cache rate
+ generated and reasoning tokens × output rate
+ tool, storage, image, audio, or search charges
Run at least three task types. A support classification job, a coding agent, and a document analysis can have different input-to-output ratios and cache hit rates. Count retries and rejected results because they still consume engineering time and may consume tokens.
Compare model coverage
Decide whether you want one model family, a catalog of open-weight models, or several families behind one account. Record the exact model IDs, lifecycle policy, context window, modalities, and current availability.
A broad catalog reduces account work only when it contains models that pass your task tests. A focused API can make sense when one model family covers the workload and its vendor-specific features matter more than portability.
Test the documented API features
Create a small compatibility matrix for the fields your application uses:
- Model discovery and stable model IDs.
- Chat completions or another required endpoint.
- Server-sent streaming and cancellation behavior.
- Tool calls, including the tool-result continuation.
- Structured output or JSON schema handling.
- Image, audio, embedding, and batch endpoints where needed.
- Error shapes, request identifiers, limits, and retry guidance.
“OpenAI-compatible” has boundaries. Test your actual SDK version and request fields. Do not assume that a compatible chat-completions endpoint includes every endpoint or vendor extension.
For services offering several model families through one API, use this production reliability checklist to evaluate complete workflows.
Review regional options and the DPA
Ask where the selected model processes a request. Separate inference processing, storage, support access, and account operations in your review. Check whether a regional option applies by default, by model, by endpoint, or through a separate commercial agreement.
Obtain the current DPA or AVV if your organization needs one. Confirm its scope, execution process, subprocessor information, and update mechanism. Your legal and privacy owners should assess the document against the intended data and workload.
Read the billing terms
Token rates do not show every cash-flow or accounting difference. Compare:
- Prepaid credits, postpaid invoices, and minimum purchases.
- Credit expiration and refund terms.
- Currency, taxes, and invoice availability.
- Spend caps, project budgets, and usage exports.
- Batch discounts, priority tiers, regional multipliers, and tool charges.
Set a spend limit and alert before a production test. Reconcile the provider dashboard with the token and request counts recorded by your application.
LLMBase publishes its EU-hosted model catalog and current prices on the inference page. If you already use an OpenAI SDK client, follow the OpenAI-compatible migration guide and keep a minimal integration test for every feature your application requires.