Nvidia logo

Nemotron 3 Ultra – Benchmarks, Pricing & Intelligence Analysis

550B

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Not available on LLMBase

Nemotron 3 Ultra is not currently available in LLMBase Chat or the LLMBase Inference API. Benchmarks, pricing, and model details remain available for research and comparison.

Input Price$0.60/1M tokens
Output Price$2.60/1M tokens
Intelligence29.6
Coding49.3

Specifications

Technical details and pricing.

ProviderNvidia
Context Window256,000 tokens
Release DateJun 4, 2026
ModalitiesText
CapabilitiesFunction Calling
AvailabilityNot available

Inference API

Power your AI projects with open-source models.

OpenAI-compatible API. Processing terms vary by selected model; review the current Privacy Policy and DPA before using sensitive data.

MiniMax

MiniMax M3

$0.40 / $1.40

per M tokens

Z.ai

GLM 5.3 Flash

$0.20 / $0.60

per M tokens

MoonshotAI

Kimi K3

$4.00 / $18.00

per M tokens

DeepSeek

DeepSeek V4 Pro

$1.80 / $3.60

per M tokens

Frequently Asked Questions

What is Nemotron 3 Ultra good for?

Use Nemotron 3 Ultra for everyday tasks like writing, summarizing, brainstorming, and getting clear explanations.

How much does Nemotron 3 Ultra cost?

Pricing is based on usage. Current rates are $0.60/1M tokens for input and $2.60/1M tokens for output.

Is Nemotron 3 Ultra available in LLMBase Chat yet?

Not yet on this page. We will add chat and inference actions automatically once an exact catalog entry exists for Nemotron 3 Ultra.

Does Nemotron 3 Ultra support images or audio?

Nemotron 3 Ultra focuses on text-based tasks.

Benchmarks and pricing use Artificial Analysis where available. Catalog specs are used as a fallback.