OpenAI-compatible Inference API

Production-ready AI inference.

Reliable performance, predictable costs, and a live model catalog. For sensitive data, review the privacy documents and confirm the terms that apply to the selected model.

Privacy Policy

Published

Model Terms

Vary by model

Data Handling

Described by service

Models & Pricing

14 Open-weight models in the current catalog

Open-weight models from the current catalog. Processing terms vary by model. Pay per use with no monthly fee.

ModelInput /1MOutput /1M
moonshotai

Kimi K3

moonshotai/kimi-k3

$4.00$18.00
z-ai

GLM 5.2

z-ai/glm-5.2

$1.40$4.40
moonshotai

Kimi K2.7 Code

moonshotai/kimi-k2.7-code

$0.95$4.00
minimax

MiniMax M3

minimax/minimax-m3

$0.40$1.40
deepseek

DeepSeek V4 Pro

deepseek/deepseek-v4-pro

$1.80$3.60
deepseek

DeepSeek V4 Flash

deepseek/deepseek-v4-flash

$0.20$0.30
qwen

Qwen3.6-35B-A3B

qwen/qwen3.6-35b-a3b

$0.20$1.20
moonshotai

Kimi K2.6

moonshotai/kimi-k2.6

$0.90$4.00
qwen

Qwen3.5 397B A17B

qwen/qwen3.5-397b-a17b

$0.60$3.60
black-forest-labs

FLUX 2 Klein

black-forest-labs/flux.2-klein-4b

$0.02 / image
qwen

Qwen3 Embedding 8B

qwen/qwen3-embedding-8b

$0.02$0.00
qwen

Qwen3-VL-30B-A3B-Instruct

qwen/qwen3-vl-30b-a3b-instruct

$0.20$0.70
openai

gpt-oss-120b

openai/gpt-oss-120b

$0.18$0.72
openai

gpt-oss-20b

openai/gpt-oss-20b

$0.07$0.20
EUProcessing location depends on the selected model and available infrastructure.
API documentation

Estimate your cost

What would your monthly volume cost?

Pick a model and a typical workload, or adjust monthly token volume, input/output ratio, and cache hit rate yourself.

$0.00

$0.00 blended / 1M tokens

Presets:

Tokens / month

Input / output ratio

85 / 15

Cache hit rate

60%

Plug and Play

Serve the latest AI models via API

We offer full compatibility with OpenAI API, allowing you to easily integrate powerful language models into your applications using OpenAI's official libraries.

Uptime

in 30 days

99.999%
Latency

on average

45ms
per Million Tokens

(*Depends on model)

$0.20
import OpenAI from 'openai';

const openai = new OpenAI({
  apiKey: process.env.LLMBASE_API_KEY,
  baseURL: 'https://api.llmbase.ai/v1'
});

const chat = await openai.chat.completions.create({
  model: "moonshotai/kimi-k3",
  messages: [{ role: "user", content: "Hello!" }],
});

Privacy & Compliance

Privacy documents and controls

Privacy documentation
Current Privacy Policy and DPA
Security measures
Documented in the current DPA
Clear data terms
Depends on the selected model and infrastructure
Data handling
Described by service in the Privacy Policy and DPA
DPA
directly in your dashboard — sign now

Processing information

Documented by service

The current DPA lists Germany and Finland for LLMBase cloud infrastructure and identifies additional subprocessors and locations by service. Model-request processing location depends on the selected model.

LLMBase Inference

Live in 60 seconds.

Get your API key and swap the base URL. Processing terms vary by selected model.

Pay-per-use · No monthly fees · Privacy Policy and DPA available

Enterprise

Custom infrastructure for your business

Dedicated GPU endpoints, agreed service targets, custom models, extended compliance documentation, and dedicated support.

  • Dedicated GPU clusters
  • Custom burst limits
  • Documented service targets
  • Custom models & fine-tuning
  • Extended compliance docs
  • Dedicated account manager