Production-ready AI inference.
Reliable performance, predictable costs, and a live model catalog. For sensitive data, review the privacy documents and confirm the terms that apply to the selected model.
Privacy Policy
Published
Model Terms
Vary by model
Data Handling
Described by service
Models & Pricing
14 Open-weight models in the current catalog
Open-weight models from the current catalog. Processing terms vary by model. Pay per use with no monthly fee.
| Model | Context | Cache /1M | Input /1M | Output /1M |
|---|---|---|---|---|
Kimi K3 moonshotai/kimi-k3 | 1M | $0.40 | $4.00 | $18.00 |
GLM 5.2 z-ai/glm-5.2 | 1M | $0.30 | $1.40 | $4.40 |
Kimi K2.7 Code moonshotai/kimi-k2.7-code | 256K | $0.20 | $0.95 | $4.00 |
MiniMax M3 minimax/minimax-m3 | 512K | $0.15 | $0.40 | $1.40 |
DeepSeek V4 Pro deepseek/deepseek-v4-pro | 1M | $0.20 | $1.80 | $3.60 |
DeepSeek V4 Flash deepseek/deepseek-v4-flash | 1M | $0.07 | $0.20 | $0.30 |
Qwen3.6-35B-A3B qwen/qwen3.6-35b-a3b | 256K | $0.10 | $0.20 | $1.20 |
Kimi K2.6 moonshotai/kimi-k2.6 | 256K | $0.20 | $0.90 | $4.00 |
Qwen3.5 397B A17B qwen/qwen3.5-397b-a17b | 256K | $0.30 | $0.60 | $3.60 |
FLUX 2 Klein black-forest-labs/flux.2-klein-4b | — | — | $0.02 / image | — |
Qwen3 Embedding 8B qwen/qwen3-embedding-8b | 31K | $0.00 | $0.02 | $0.00 |
Qwen3-VL-30B-A3B-Instruct qwen/qwen3-vl-30b-a3b-instruct | 256K | $0.00 | $0.20 | $0.70 |
gpt-oss-120b openai/gpt-oss-120b | 128K | $0.09 | $0.18 | $0.72 |
gpt-oss-20b openai/gpt-oss-20b | 128K | $0.00 | $0.07 | $0.20 |
Estimate your cost
What would your monthly volume cost?
Pick a model and a typical workload, or adjust monthly token volume, input/output ratio, and cache hit rate yourself.
$0.00
$0.00 blended / 1M tokens
Tokens / month
Input / output ratio
85 / 15Cache hit rate
60%Plug and Play
Serve the latest AI models via API
We offer full compatibility with OpenAI API, allowing you to easily integrate powerful language models into your applications using OpenAI's official libraries.
- Uptime
in 30 days
- 99.999%
- Latency
on average
- 45ms
- per Million Tokens
(*Depends on model)
- $0.20
import OpenAI from 'openai';
const openai = new OpenAI({
apiKey: process.env.LLMBASE_API_KEY,
baseURL: 'https://api.llmbase.ai/v1'
});
const chat = await openai.chat.completions.create({
model: "moonshotai/kimi-k3",
messages: [{ role: "user", content: "Hello!" }],
});Privacy & Compliance
Privacy documents and controls
- Privacy documentation
- Current Privacy Policy and DPA
- Security measures
- Documented in the current DPA
- Clear data terms
- Depends on the selected model and infrastructure
- Data handling
- Described by service in the Privacy Policy and DPA
- DPA
- directly in your dashboard — sign now
Processing information
Documented by service
The current DPA lists Germany and Finland for LLMBase cloud infrastructure and identifies additional subprocessors and locations by service. Model-request processing location depends on the selected model.
LLMBase Inference
Live in 60 seconds.
Get your API key and swap the base URL. Processing terms vary by selected model.
Pay-per-use · No monthly fees · Privacy Policy and DPA available
Enterprise
Custom infrastructure for your business
Dedicated GPU endpoints, agreed service targets, custom models, extended compliance documentation, and dedicated support.
- Dedicated GPU clusters
- Custom burst limits
- Documented service targets
- Custom models & fine-tuning
- Extended compliance docs
- Dedicated account manager
