back

OLMo 2 32B

Apache 2.0

Allen AI · 32B · Dense

Fully open research model by Allen AI

3.7K downloads 148 likes 2025-03 4K context

Use Cases

chat reasoning

Quantization Options

Quant Bits VRAM Quality Status
Q2_K 2 10.7 GB low
Q3_K_M 3 14.8 GB moderate
Q4_K_M 4 16.9 GB good
Q5_K_M 5 21 GB good
Q6_K 6 25.1 GB excellent
Q8_0 8 33.3 GB excellent
F16 16 66.1 GB lossless

About this model

HF model card OLMo

Allen AI

32B

Dense model for chat, reasoning workloads.

Context

4K

Q4 VRAM

~16.9 GB

Dense

OLMo 2 32B on your hardware

Fully open research model by Allen AI. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.

Model shape

32B total parameters. Dense models use the whole network for each token.

Local fit target

Start with Q4_K_M: about 15.6 GB on disk and 16.9 GB VRAM before extra context and runtime overhead.

Best use cases

chat, reasoning

License and source

Listed as Apache 2.0 from huggingface.co/allenai/OLMo-2-0325-32B-Instruct.

Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.

Open HF repo