Qwen logo

Qwen3.5-Flash

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

Input Price$0.10/1M tokens
Output Price$0.80/1M tokens
Intelligence19.0
CodingN/A

Specifications

Technical details and pricing.

ProviderQwen
Context Window1,000,000 tokens
Release DateMar 30, 2026
ModalitiesText, Image β†’ Text
CapabilitiesFunction Calling, Vision
AvailabilityChat only

Benchmarks

7 benchmark scores from Artificial Analysis.

GPQA74.2%
HLE7.1%
SciCode25.5%
LCR44.0%
IFBench38.0%
Tau284.5%
TerminalBench Hard8.3%

Composite Indices

Higher is better; speed and price are normalized

Standard Benchmarks

Only benchmarks with data are shown

EU-Hosted Inference API

Power your AI projects with the best open-source models.

Drop-in OpenAI-compatible API. No data leaves Europe.

Explore Inference API

MiniMax

MiniMax M3

$0.40 / $1.40

per M tokens

Z.ai

GLM 5.2

$1.40 / $4.40

per M tokens

MoonshotAI

Kimi K2.7 Code

$0.95 / $4.00

per M tokens

DeepSeek

DeepSeek V4 Pro

$1.80 / $3.60

per M tokens

Frequently Asked Questions

What is Qwen3.5-Flash good for?

Use Qwen3.5-Flash for everyday tasks like writing, summarizing, brainstorming, and getting clear explanations.

How much does Qwen3.5-Flash cost?

Pricing is based on usage. Current rates are $0.10/1M tokens for input and $0.80/1M tokens for output.

Can I use Qwen3.5-Flash in LLMBase Chat?

Yes. Use the chat action on this page to open Qwen3.5-Flash in LLMBase Chat. Access still follows the current plan and model tier.

Does Qwen3.5-Flash support images or audio?

Qwen3.5-Flash can understand images.

Benchmarks and pricing use Artificial Analysis where available. Catalog specs are used as a fallback.