Available AI Models

AI Models on LLMBase

Compare models available in LLMBase Chat, through the Inference API, or both.

EU-hosted Open Models

Every available open-weight model, newest first.

12 models
MoonshotAI logo

Kimi K3

EU-Hosted

MoonshotAI

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Context1.0M
Speed41 tok/s
InputText, Image
OutputText
ReasoningYes
AvailabilityChat, Inference API
Z.ai logo

GLM 5.2

EU-Hosted

Z.ai

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Context1.0M
Speed96 tok/s
InputText
OutputText
ReasoningYes
AvailabilityChat, Inference API
MoonshotAI logo

MoonshotAI

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Context262K
Speed46 tok/s
InputText
OutputText
ReasoningNo
AvailabilityChat, Inference API
MiniMax logo

M3

EU-Hosted

MiniMax

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Context524K
Speed108 tok/s
InputText, Image
OutputText
ReasoningYes
AvailabilityChat, Inference API
DeepSeek logo

V4 Flash 0423

EU-Hosted

DeepSeek

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Context1.0M
SpeedN/A
InputText
OutputText
ReasoningYes
AvailabilityChat, Inference API
DeepSeek logo

V4 Pro 0423

EU-Hosted

DeepSeek

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Context1.0M
Speed68 tok/s
InputText
OutputText
ReasoningYes
AvailabilityChat, Inference API
MoonshotAI logo

Kimi K2.6

EU-Hosted

MoonshotAI

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Context262K
Speed129 tok/s
InputText, Image
OutputText
ReasoningYes
AvailabilityChat, Inference API
Qwen logo

Qwen

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

Context262K
Speed165 tok/s
InputText, Image
OutputText
ReasoningNo
AvailabilityChat, Inference API
Qwen logo

Qwen

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

Context262K
Speed90 tok/s
InputText, Image
OutputText
ReasoningNo
AvailabilityChat, Inference API
Black Forest Labs logo

Black Forest Labs

Fast Black Forest Labs image generation and editing model.

ContextN/A
SpeedN/A
InputText, Image
OutputN/A
ReasoningNo
Availability
OpenAI logo

gpt-oss-120b

EU-Hosted

OpenAI

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

Context131K
Speed156 tok/s
InputText
OutputText
ReasoningNo
AvailabilityChat, Inference API
OpenAI logo

gpt-oss-20b

EU-Hosted

OpenAI

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

Context131K
Speed148 tok/s
InputText
OutputText
ReasoningNo
AvailabilityChat, Inference API

Proprietary Models

Every available proprietary model, newest first.

41 models
Qwen logo

Qwen3.8 Flash

NEWProprietary

Qwen

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Context1.0M
Speed190 tok/s
InputText, Image, Video
OutputText
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.8 27B

Proprietary

Qwen

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Context1.0M
Speed83 tok/s
InputText, Image, Video
OutputText
ReasoningNo
AvailabilityChat only
Google logo

Gemini 3.7 Flash

Proprietary

Google

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Context1.0M
Speed322 tok/s
InputText, Image, Video, File, Audio
OutputText
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Context1.0M
Speed24 tok/s
InputText
OutputText
ReasoningNo
AvailabilityChat only
X Ai logo

X Ai

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Context500K
Speed60 tok/s
InputText, Image, File
OutputText
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.8 Max

Proprietary

Qwen

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...

Context1.0M
Speed60 tok/s
InputText, Image, Video
OutputText
ReasoningNo
AvailabilityChat only
Anthropic logo

Anthropic

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Context1.0M
SpeedN/A
InputText, Image, File
OutputText
ReasoningNo
AvailabilityChat only
Anthropic logo

Claude Opus 5

Proprietary

Anthropic

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Context1.0M
Speed126 tok/s
InputText, Image, File
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.6 Luna Pro

Proprietary

Openai

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context1.1M
SpeedN/A
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

Openai

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context1.1M
SpeedN/A
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.6 Sol Pro

Proprietary

Openai

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context1.1M
SpeedN/A
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.6 Luna

Proprietary

Openai

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Context1.1M
Speed160 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.6 Sol

Proprietary

Openai

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Context1.1M
Speed75 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.6 Terra

Proprietary

Openai

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Context1.1M
Speed108 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Context66K
SpeedN/A
InputImage, Text
OutputImage, Text
ReasoningNo
AvailabilityChat only
Anthropic logo

Claude Sonnet 5

Proprietary

Anthropic

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Context1.0M
Speed86 tok/s
InputText, Image, File
OutputText
ReasoningNo
AvailabilityChat only
Anthropic logo

Claude Fable 5

Proprietary

Anthropic

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Context1.0M
Speed65 tok/s
InputText, Image, File
OutputText
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.7 Plus

Proprietary

Qwen

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Context1.0M
Speed56 tok/s
InputText, Image
OutputText
ReasoningYes
AvailabilityChat only

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

Context131K
Speed195 tok/s
InputImage, Text
OutputImage, Text
ReasoningNo
AvailabilityChat only
Google logo

Google

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Context1.0M
Speed391 tok/s
InputText, Image, Video, File, Audio
OutputText
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.7 Max

Proprietary

Qwen

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Context1.0M
Speed199 tok/s
InputText
OutputText
ReasoningYes
AvailabilityChat only
Mistralai logo

Mistralai

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Context262K
Speed142 tok/s
InputText, Image, File
OutputText
ReasoningNo
AvailabilityChat only

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

Context1.0M
SpeedN/A
InputText, Image, Video
OutputText
ReasoningYes
AvailabilityChat only
Qwen logo

Qwen3.6 Flash

Proprietary

Qwen

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

Context1.0M
SpeedN/A
InputText, Image, Video
OutputText
ReasoningYes
AvailabilityChat only
Openai logo

GPT-5.5

Proprietary

Openai

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Context1.1M
Speed81 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.5 Pro

Proprietary

Openai

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Context1.1M
Speed148 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.6 27B

Proprietary

Qwen

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

Context262K
Speed54 tok/s
InputText, Image, Video
OutputText
ReasoningYes
AvailabilityChat only
Qwen logo

Qwen

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Context262K
Speed35 tok/s
InputText
OutputText
ReasoningYes
AvailabilityChat only
Google logo

Gemma 4 26B A4B

Proprietary

Google

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Context262K
Speed113 tok/s
InputImage, Text, Video
OutputText
ReasoningYes
AvailabilityChat only
Google logo

Gemma 4 31B

Proprietary

Google

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Context262K
Speed81 tok/s
InputImage, Text, Video
OutputText
ReasoningYes
AvailabilityChat only
Qwen logo

Qwen3.6 Plus

Proprietary

Qwen

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Context1.0M
Speed55 tok/s
InputText, Image, Video
OutputText
ReasoningYes
AvailabilityChat only
Google logo

Google

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

Context1.0M
SpeedN/A
InputText, Image
OutputText, Audio
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.7 Flash

Proprietary

Qwen

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Context1.0M
Speed250 tok/s
InputText, Image, Video
OutputText
ReasoningNo
AvailabilityChat only
Minimax logo

M2.7

Proprietary

Minimax

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

Context205K
Speed95 tok/s
InputText
OutputText
ReasoningYes
AvailabilityChat only
Openai logo

GPT-5.4 Mini

Proprietary

Openai

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

Context400K
Speed173 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.4 Nano

Proprietary

Openai

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

Context400K
Speed172 tok/s
InputFile, Image, Text
OutputText
ReasoningNo
AvailabilityChat only
Mistralai logo

Mistralai

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

Context262K
Speed153 tok/s
InputText, Image
OutputText
ReasoningNo
AvailabilityChat only
Openai logo

GPT-5.4 Image 2

Proprietary

Openai

It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

Context272K
Speed130 tok/s
InputImage, Text, File
OutputImage, Text
ReasoningNo
AvailabilityChat only
Qwen logo

Qwen3.5-9B

Proprietary

Qwen

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

Context262K
Speed94 tok/s
InputText, Image, Video
OutputText
ReasoningYes
AvailabilityChat only
Google logo

Google

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

Context1.0M
Speed115 tok/s
InputText, Image
OutputText, Audio
ReasoningNo
AvailabilityChat only

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Context131K
Speed124 tok/s
InputImage, Text
OutputImage, Text
ReasoningYes
AvailabilityChat only

Secure AI Chat made in Europe

Use Claude, ChatGPT, Gemini alongside with EU-Hosted Models like Deepseek, Qwen & Kimi.

EU-hosted inference

Servers in Germany & Finland. Designed to meet strict GDPR and ISO 27001 compliance requirements.

How to Choose the Right AI Model

A practical guide to picking the best LLM for your use case.

Match the model to the task

General-purpose models handle most tasks well. For specialized work, coding models and math models often outperform generalists on their respective benchmarks while costing less per token.

Consider context window size

If you work with long documents, codebases, or multi-turn conversations, context window matters. Models range from 8K to over 1M tokens. Larger windows let you process entire books or repositories in a single prompt, but they increase cost and latency.

Balance cost, speed, and quality

Frontier models deliver the highest benchmark scores but cost more per token and respond slower. Faster smaller models can handle routine tasks at a fraction of the cost with lower latency, which is often better for high-volume applications.

Open source vs. proprietary

Open-source models let you self-host, fine-tune, and inspect weights. Proprietary models often lead on benchmarks and offer managed APIs with built-in safety features. Many teams use both: proprietary for peak performance, open source for cost control and customization.

Check for multimodal capabilities

Some models accept images, audio, or files alongside text. If your workflow involves analyzing screenshots, diagrams, or audio transcriptions, filter for models with vision or audio input support. Models with structured output and function calling are essential for building agents and tool-using applications.

Use benchmarks as a starting point

Scores like GPQA, MMLU Pro, and HLE measure academic knowledge and reasoning. LiveCodeBench and SciCode test practical coding ability. MATH 500 and AIME evaluate mathematical problem-solving. No single benchmark tells the full story - compare scores across categories relevant to your use case, then test with your own prompts.

Model catalog, pricing, speed, and benchmark scores are updated regularly. Chat and API access follows each model, plan and tier.