back

Gemma 4 26B-A4B IT

Gemma

Google · 27B (4B active) · Mixture of Experts

Gemma 4 MoE instruct model (official)

0 downloads 0 likes 2026-04 256K context

Use Cases

chat vision reasoning

Mixture of Experts

Total experts: 26
Active experts: 4
Active params: 4.0B

Quantization Options

Quant Bits VRAM Quality Status
Q4_K_M 4 14.3 GB good
Q6_K 6 21.2 GB excellent
Q8_0 8 28.2 GB excellent

About this model

HF model card Gemma

Google

27B

Mixture of Experts model for chat, vision, reasoning workloads.

Context

256K

Q4 VRAM

~14.3 GB

Mixture of Experts Vision

Gemma 4 26B-A4B IT on your hardware

Gemma 4 MoE instruct model (official). This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.

Model shape

27B total parameters with 4B active. 4 of 26 experts are active per token.

Local fit target

Start with Q4_K_M: about 15.6 GB on disk and 14.3 GB VRAM before extra context and runtime overhead.

Best use cases

chat, vision, reasoning

License and source

Listed as Gemma from huggingface.co/lmstudio-community/gemma-4-26B-A4B-it-GGUF.

Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.

Open HF repo