Guides

LLM API providers and pricing in 2026: a practical comparison

Compare representative public LLM API prices, caching, model coverage, API features, regional options, DPA availability, and billing terms.

LLM API providers and pricing in 2026: a practical comparison

LLM API price tables show a rate, not the cost of a completed task. Compare input, cached input, output, tool charges, model behavior, and billing terms with the same workload.

This guide compares public products that developers can buy directly from each company. Prices below are representative public list prices checked on 31 August 2026, shown in USD per one million text tokens unless noted. Vendors can change prices and model availability. Open the linked primary source before making a purchase or production decision.

Representative public prices

Public API and modelInputCached input or cache readOutputPrimary source checked 31 Aug 2026
LLMBase, DeepSeek V4 Flash$0.20$0.07$0.30LLMBase live model catalog
OpenAI, GPT-5.6 Luna$0.20$0.02$1.20OpenAI model documentation
Anthropic, Sonnet 5$2.00$0.20 cache read$10.00Anthropic API pricing
Google, Gemini 3.5 Flash$1.50$0.15 plus cache storage$9.00Gemini API pricing
Mistral, Mistral Large$0.50Up to 90% lower for cached input$1.50Mistral pricing

These models do not have equal quality, speed, context, or tool support. The table illustrates pricing structures. It does not rank model capability.

Caching also means different things across APIs. Check whether a vendor charges for cache writes, reads, storage duration, or automatic reuse. A low cache-read rate has little value when your prompts rarely share a stable prefix.

Calculate the cost of your workload

Use measured token counts from a representative request. A basic monthly estimate is:

uncached input tokens × input rate
+ cached tokens × cache rate
+ generated and reasoning tokens × output rate
+ tool, storage, image, audio, or search charges

Run at least three task types. A support classification job, a coding agent, and a document analysis can have different input-to-output ratios and cache hit rates. Count retries and rejected results because they still consume engineering time and may consume tokens.

Compare model coverage

Decide whether you want one model family, a catalog of open-weight models, or several families behind one account. Record the exact model IDs, lifecycle policy, context window, modalities, and current availability.

A broad catalog reduces account work only when it contains models that pass your task tests. A focused API can make sense when one model family covers the workload and its vendor-specific features matter more than portability.

Test the documented API features

Create a small compatibility matrix for the fields your application uses:

  • Model discovery and stable model IDs.
  • Chat completions or another required endpoint.
  • Server-sent streaming and cancellation behavior.
  • Tool calls, including the tool-result continuation.
  • Structured output or JSON schema handling.
  • Image, audio, embedding, and batch endpoints where needed.
  • Error shapes, request identifiers, limits, and retry guidance.

“OpenAI-compatible” has boundaries. Test your actual SDK version and request fields. Do not assume that a compatible chat-completions endpoint includes every endpoint or vendor extension.

For services offering several model families through one API, use this production reliability checklist to evaluate complete workflows.

Review regional options and the DPA

Ask where the selected model processes a request. Separate inference processing, storage, support access, and account operations in your review. Check whether a regional option applies by default, by model, by endpoint, or through a separate commercial agreement.

Obtain the current DPA or AVV if your organization needs one. Confirm its scope, execution process, subprocessor information, and update mechanism. Your legal and privacy owners should assess the document against the intended data and workload.

Read the billing terms

Token rates do not show every cash-flow or accounting difference. Compare:

  • Prepaid credits, postpaid invoices, and minimum purchases.
  • Credit expiration and refund terms.
  • Currency, taxes, and invoice availability.
  • Spend caps, project budgets, and usage exports.
  • Batch discounts, priority tiers, regional multipliers, and tool charges.

Set a spend limit and alert before a production test. Reconcile the provider dashboard with the token and request counts recorded by your application.

LLMBase publishes its EU-hosted model catalog and current prices on the inference page. If you already use an OpenAI SDK client, follow the OpenAI-compatible migration guide and keep a minimal integration test for every feature your application requires.

All guides