Verfügbare KI-Modelle

KI-Modelle auf LLMBase

Vergleiche Modelle, die direkt in LLMBase Chat, über die Inference API oder auf beiden Wegen verfügbar sind.

EU-gehostete offene Modelle

Alle verfügbaren Open-Weight-Modelle, neueste zuerst.

12 Modelle
MoonshotAI logo

Kimi K3

EU-gehostet

MoonshotAI

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Kontext1.0M
Geschwindigkeit41 tok/s
EingabeText, Image
AusgabeText
ReasoningJa
VerfügbarkeitChat, Inference API
Z.ai logo

GLM 5.2

EU-gehostet

Z.ai

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Kontext1.0M
Geschwindigkeit96 tok/s
EingabeText
AusgabeText
ReasoningJa
VerfügbarkeitChat, Inference API
MoonshotAI logo

Kimi K2.7 Code

EU-gehostet

MoonshotAI

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Kontext262K
Geschwindigkeit45 tok/s
EingabeText
AusgabeText
ReasoningNein
VerfügbarkeitChat, Inference API
MiniMax logo

M3

EU-gehostet

MiniMax

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Kontext524K
Geschwindigkeit142 tok/s
EingabeText, Image
AusgabeText
ReasoningJa
VerfügbarkeitChat, Inference API
DeepSeek logo

V4 Flash 0423

EU-gehostet

DeepSeek

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Kontext1.0M
GeschwindigkeitN/A
EingabeText
AusgabeText
ReasoningJa
VerfügbarkeitChat, Inference API
DeepSeek logo

V4 Pro 0423

EU-gehostet

DeepSeek

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Kontext1.0M
Geschwindigkeit66 tok/s
EingabeText
AusgabeText
ReasoningJa
VerfügbarkeitChat, Inference API
MoonshotAI logo

Kimi K2.6

EU-gehostet

MoonshotAI

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Kontext262K
Geschwindigkeit129 tok/s
EingabeText, Image
AusgabeText
ReasoningJa
VerfügbarkeitChat, Inference API
Qwen logo

Qwen3.6 35B A3B

EU-gehostet

Qwen

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

Kontext262K
Geschwindigkeit165 tok/s
EingabeText, Image
AusgabeText
ReasoningNein
VerfügbarkeitChat, Inference API
Qwen logo

Qwen

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

Kontext262K
Geschwindigkeit80 tok/s
EingabeText, Image
AusgabeText
ReasoningNein
VerfügbarkeitChat, Inference API
Black Forest Labs logo

Black Forest Labs

Fast Black Forest Labs image generation and editing model.

KontextN/A
GeschwindigkeitN/A
EingabeText, Image
AusgabeN/A
ReasoningNein
Verfügbarkeit
OpenAI logo

gpt-oss-120b

EU-gehostet

OpenAI

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

Kontext131K
Geschwindigkeit156 tok/s
EingabeText
AusgabeText
ReasoningNein
VerfügbarkeitChat, Inference API
OpenAI logo

gpt-oss-20b

EU-gehostet

OpenAI

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

Kontext131K
Geschwindigkeit145 tok/s
EingabeText
AusgabeText
ReasoningNein
VerfügbarkeitChat, Inference API

Proprietäre Modelle

Alle verfügbaren proprietären Modelle, neueste zuerst.

41 Modelle
Qwen logo

Qwen3.8 Flash

NEWProprietär

Qwen

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Kontext1.0M
Geschwindigkeit190 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.8 27B

Proprietär

Qwen

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Kontext1.0M
Geschwindigkeit82 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Google logo

Gemini 3.7 Flash

Proprietär

Google

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Kontext1.0M
Geschwindigkeit322 tok/s
EingabeText, Image, Video, File, Audio
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Kontext1.0M
Geschwindigkeit39 tok/s
EingabeText
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
X Ai logo

X Ai

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Kontext500K
Geschwindigkeit78 tok/s
EingabeText, Image, File
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.8 Max

Proprietär

Qwen

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...

Kontext1.0M
Geschwindigkeit59 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Anthropic logo

Anthropic

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Kontext1.0M
GeschwindigkeitN/A
EingabeText, Image, File
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Anthropic logo

Claude Opus 5

Proprietär

Anthropic

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Kontext1.0M
Geschwindigkeit126 tok/s
EingabeText, Image, File
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.6 Luna Pro

Proprietär

Openai

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Kontext1.1M
GeschwindigkeitN/A
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

Openai

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Kontext1.1M
GeschwindigkeitN/A
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.6 Sol Pro

Proprietär

Openai

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Kontext1.1M
GeschwindigkeitN/A
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.6 Luna

Proprietär

Openai

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Kontext1.1M
Geschwindigkeit179 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.6 Sol

Proprietär

Openai

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Kontext1.1M
Geschwindigkeit78 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.6 Terra

Proprietär

Openai

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Kontext1.1M
Geschwindigkeit108 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Kontext66K
GeschwindigkeitN/A
EingabeImage, Text
AusgabeImage, Text
ReasoningNein
VerfügbarkeitNur Chat
Anthropic logo

Claude Sonnet 5

Proprietär

Anthropic

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Kontext1.0M
Geschwindigkeit71 tok/s
EingabeText, Image, File
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Anthropic logo

Claude Fable 5

Proprietär

Anthropic

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Kontext1.0M
Geschwindigkeit65 tok/s
EingabeText, Image, File
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.7 Plus

Proprietär

Qwen

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Kontext1.0M
Geschwindigkeit56 tok/s
EingabeText, Image
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

Kontext131K
Geschwindigkeit183 tok/s
EingabeImage, Text
AusgabeImage, Text
ReasoningNein
VerfügbarkeitNur Chat
Google logo

Google

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Kontext1.0M
Geschwindigkeit366 tok/s
EingabeText, Image, Video, File, Audio
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.7 Max

Proprietär

Qwen

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Kontext1.0M
Geschwindigkeit205 tok/s
EingabeText
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Mistralai logo

Mistralai

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Kontext262K
Geschwindigkeit148 tok/s
EingabeText, Image, File
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

Kontext1.0M
GeschwindigkeitN/A
EingabeText, Image, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Qwen logo

Qwen3.6 Flash

Proprietär

Qwen

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

Kontext1.0M
GeschwindigkeitN/A
EingabeText, Image, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Openai logo

GPT-5.5

Proprietär

Openai

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Kontext1.1M
Geschwindigkeit94 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.5 Pro

Proprietär

Openai

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Kontext1.1M
Geschwindigkeit148 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.6 27B

Proprietär

Qwen

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

Kontext262K
Geschwindigkeit55 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Qwen logo

Qwen

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Kontext262K
Geschwindigkeit58 tok/s
EingabeText
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Google logo

Gemma 4 26B A4B

Proprietär

Google

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Kontext262K
Geschwindigkeit111 tok/s
EingabeImage, Text, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Google logo

Gemma 4 31B

Proprietär

Google

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Kontext262K
Geschwindigkeit81 tok/s
EingabeImage, Text, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Qwen logo

Qwen3.6 Plus

Proprietär

Qwen

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Kontext1.0M
Geschwindigkeit55 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Google logo

Google

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

Kontext1.0M
GeschwindigkeitN/A
EingabeText, Image
AusgabeText, Audio
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.7 Flash

Proprietär

Qwen

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Kontext1.0M
Geschwindigkeit227 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Minimax logo

M2.7

Proprietär

Minimax

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

Kontext205K
Geschwindigkeit98 tok/s
EingabeText
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Openai logo

GPT-5.4 Mini

Proprietär

Openai

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

Kontext400K
Geschwindigkeit197 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.4 Nano

Proprietär

Openai

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

Kontext400K
Geschwindigkeit186 tok/s
EingabeFile, Image, Text
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Mistralai logo

Mistralai

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

Kontext262K
Geschwindigkeit171 tok/s
EingabeText, Image
AusgabeText
ReasoningNein
VerfügbarkeitNur Chat
Openai logo

GPT-5.4 Image 2

Proprietär

Openai

It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

Kontext272K
Geschwindigkeit130 tok/s
EingabeImage, Text, File
AusgabeImage, Text
ReasoningNein
VerfügbarkeitNur Chat
Qwen logo

Qwen3.5-9B

Proprietär

Qwen

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

Kontext262K
Geschwindigkeit96 tok/s
EingabeText, Image, Video
AusgabeText
ReasoningJa
VerfügbarkeitNur Chat
Google logo

Google

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

Kontext1.0M
Geschwindigkeit115 tok/s
EingabeText, Image
AusgabeText, Audio
ReasoningNein
VerfügbarkeitNur Chat

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Kontext131K
Geschwindigkeit128 tok/s
EingabeImage, Text
AusgabeImage, Text
ReasoningJa
VerfügbarkeitNur Chat

Sicherer KI-Chat aus Europa

Nutze Claude, ChatGPT und Gemini zusammen mit EU-gehosteten Modellen wie Deepseek, Qwen und Kimi.

EU-gehostete Inferenz

Server in Deutschland und Finnland. Entwickelt fuer strenge GDPR- und ISO-27001-Compliance-Anforderungen.

So findest du das richtige KI-Modell

Ein praktischer Leitfaden für die Modellwahl nach Use Case.

Modell an Aufgabe anpassen

Allzweckmodelle funktionieren für viele Aufgaben gut. Für Spezialfälle sind Coding-Modelle und Mathe-Modelle oft präziser und günstiger pro Token.

Kontextfenster berücksichtigen

Für lange Dokumente, Codebasen oder lange Konversationen ist die Kontextgröße entscheidend. Modelle reichen von 8K bis über 1M Token. Größere Fenster erlauben mehr Input, erhöhen aber häufig Kosten und Latenz.

Kosten, Geschwindigkeit und Qualität balancieren

Frontier-Modelle liefern Top-Benchmarkwerte, sind aber teurer und oft langsamer. Schnellere, kleinere Modelle können Routineaufgaben bei hoher Last oft günstiger und mit niedrigerer Latenz erledigen.

Open Source vs. proprietär

Open-Source-Modelle bieten Self-Hosting, Anpassbarkeit und prüfbare Gewichte. Proprietäre Modelle führen häufig bei Benchmarks und bieten verwaltete APIs mit integrierten Sicherheitsfunktionen. Viele Teams kombinieren beide Ansätze.

Multimodalität prüfen

Einige Modelle verarbeiten neben Text auch Bilder, Audio oder Dateien. Für Workflows mit Screenshots, Diagrammen oder Audio-Transkripten sind Vision-/Audio-Inputs wichtig. Strukturierte Outputs und Function Calling sind zentral für Agenten.

Benchmarks als Ausgangspunkt nutzen

GPQA, MMLU Pro und HLE messen Wissen und Reasoning. LiveCodeBench und SciCode testen Coding-Praxis. MATH 500 und AIME bewerten mathematisches Lösen. Vergleiche relevante Kategorien und teste zusätzlich mit eigenen Prompts.

Modellkatalog, Preise, Geschwindigkeit und Benchmark-Scores werden regelmäßig aktualisiert. Chat- und API-Zugriff folgen dem jeweiligen Modell, Tarif und Tier.