back

DiffusionGemma 26B-A4B IT

Apache 2.0

Google · 25.2B (3.8B active) · Mixture of Experts

Discrete-diffusion multimodal Gemma model that denoises token blocks for faster generation

0 downloads 0 likes 2026-06 256K context

Use Cases

chat vision reasoning

Mixture of Experts

Total experts: 128
Active experts: 8
Active params: 3.8B

Quantization Options

Quant Bits VRAM Quality Status
Q2_K 2 8.6 GB low
Q3_K_M 3 11.8 GB moderate
Q4_K_M 4 13.4 GB good
Q5_K_M 5 16.6 GB good
Q6_K 6 19.9 GB excellent
Q8_0 8 26.3 GB excellent
F16 16 52.1 GB lossless

About this model

HF model card Gemma
DiffusionGemma 26B-A4B IT source artwork

Google

25.2B

Mixture of Experts model for chat, vision, reasoning workloads.

Context

256K

Q4 VRAM

~13.4 GB

Image saved from public model sources

Mixture of Experts Thinking Tool use Vision

DiffusionGemma 26B-A4B IT on your hardware

Discrete-diffusion multimodal Gemma model that denoises token blocks for faster generation. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.

Model shape

25.2B total parameters with 3.8B active. 8 of 128 experts are active per token.

Local fit target

Start with Q4_K_M: about 12.3 GB on disk and 13.4 GB VRAM before extra context and runtime overhead.

Best use cases

chat, vision, reasoning

License and source

Listed as Apache 2.0 from huggingface.co/google/diffusiongemma-26B-A4B-it.

Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.

Open HF repo

Benchmark snapshot

Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.

Source
8 public scores google/diffusiongemma-26B-A4B-it Hugging Face model card

Knowledge

MMLU Pro

77.6%

Math

AIME 2026

69.1%

Coding

LiveCodeBench v6

69.1%

Reasoning

GPQA Diamond

73.2%

Vision

MMMU Pro

54.3%

Document vision

OmniDocBench 1.5

0.319

Edit distance, lower is better

Long context

MRCR v2 128K

32.0%

8-needle average

Knowledge

HLE with search

11.9%