Nvidia
NVIDIA: Nemotron 3 Ultra (free)
Nvidia
Input / M
Routed
Output / M
Context
Max Output
66K
Usage
Unranked
Quality
Unscored
Providers
1
Best Uptime 30m
97.38%
Latency (p50)
28785 ms
Throughput
9.0 tok/s
Released
June 4, 2026
Modalities
text
About
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it.
Capabilities
Vision
Tools
Reasoning
Streaming
JSON Mode
Host offerings
Per-host pricing, uptime, latency
No host data yet — will populate on next sync.
Price History
0 recorded changes
No price changes recorded yet.
Activity Timeline
- No events yet.
Versus
Nearest peers by price, context, and capabilities. One click to compare.
No comparable models found.
Show this WM Score
Free live badge for your README, docs or site
Live badge — it always shows the current WM Score. Free to use anywhere, no key needed.
Markdown
[](https://whatmodel.app/models/0c3c870a-ea86-4c66-8d15-0215cf000747)HTML
<a href="https://whatmodel.app/models/0c3c870a-ea86-4c66-8d15-0215cf000747"><img src="https://whatmodel.app/api/public/embed/score/0c3c870a-ea86-4c66-8d15-0215cf000747" alt="NVIDIA: Nemotron 3 Ultra (free) WM Score" height="26" /></a>Image URL
https://whatmodel.app/api/public/embed/score/0c3c870a-ea86-4c66-8d15-0215cf000747First seen 44d ago · Last verified 6m ago
