NVIDIA · 550B (55B active) · Mixture of Experts
NVIDIA frontier hybrid LatentMoE model for agentic reasoning, tool use, RAG, and multilingual long-context work
Use Cases
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 176.6 GB | low | — |
| Q3_K_M | 3 | 247 GB | moderate | — |
| Q4_K_M | 4 | 282.2 GB | good | — |
| Q5_K_M | 5 | 352.7 GB | good | — |
| Q6_K | 6 | 423.1 GB | excellent | — |
| Q8_0 | 8 | 564 GB | excellent | — |
| F16 | 16 | 1127.4 GB | lossless | — |
About this model
NVIDIA
Mixture of Experts model for chat, reasoning, code workloads.
Context
1024K
Q4 VRAM
~282.2 GB
Image saved from public model sources
Nemotron 3 Ultra 550B-A55B on your hardware
NVIDIA frontier hybrid LatentMoE model for agentic reasoning, tool use, RAG, and multilingual long-context work. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.
550B total parameters with 55B active. Dense models use the whole network for each token.
Start with Q4_K_M: about 268.9 GB on disk and 282.2 GB VRAM before extra context and runtime overhead.
chat, reasoning, code, rag
Listed as OpenMDW 1.1 from huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16.
Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.
Open HF repoBenchmark snapshot
Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.
Coding agent
SWE-Bench Verified
70.7
Coding agent
Terminal Bench 2.1
56.4
Coding
LiveCodeBench v6
89.0
Reasoning
GPQA
87.0
No tools
Knowledge
MMLU-Pro
86.8
Long context
RULER 1M
94.7
Agentic
TauBench V3 Average
70.9
Multilingual
MMLU-ProX
83.0