Google · 25.2B (3.8B active) · Mixture of Experts
Discrete-diffusion multimodal Gemma model that denoises token blocks for faster generation
Use Cases
Mixture of Experts
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 8.6 GB | low | — |
| Q3_K_M | 3 | 11.8 GB | moderate | — |
| Q4_K_M | 4 | 13.4 GB | good | — |
| Q5_K_M | 5 | 16.6 GB | good | — |
| Q6_K | 6 | 19.9 GB | excellent | — |
| Q8_0 | 8 | 26.3 GB | excellent | — |
| F16 | 16 | 52.1 GB | lossless | — |
About this model
Mixture of Experts model for chat, vision, reasoning workloads.
Context
256K
Q4 VRAM
~13.4 GB
Image saved from public model sources
DiffusionGemma 26B-A4B IT on your hardware
Discrete-diffusion multimodal Gemma model that denoises token blocks for faster generation. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.
25.2B total parameters with 3.8B active. 8 of 128 experts are active per token.
Start with Q4_K_M: about 12.3 GB on disk and 13.4 GB VRAM before extra context and runtime overhead.
chat, vision, reasoning
Listed as Apache 2.0 from huggingface.co/google/diffusiongemma-26B-A4B-it.
Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.
Open HF repoBenchmark snapshot
Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.
Knowledge
MMLU Pro
77.6%
Math
AIME 2026
69.1%
Coding
LiveCodeBench v6
69.1%
Reasoning
GPQA Diamond
73.2%
Vision
MMMU Pro
54.3%
Document vision
OmniDocBench 1.5
0.319
Edit distance, lower is better
Long context
MRCR v2 128K
32.0%
8-needle average
Knowledge
HLE with search
11.9%