Mistral AI · 128B · Dense
Dense flagship Mistral model for instruction following, reasoning, coding agents, and multimodal work
Use Cases
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 41.5 GB | low | — |
| Q3_K_M | 3 | 57.9 GB | moderate | — |
| Q4_K_M | 4 | 66.1 GB | good | — |
| Q5_K_M | 5 | 82.5 GB | good | — |
| Q6_K | 6 | 98.8 GB | excellent | — |
| Q8_0 | 8 | 131.6 GB | excellent | — |
| F16 | 16 | 262.8 GB | lossless | — |
About this model
Mistral AI
Dense model for chat, vision, reasoning workloads.
Context
256K
Q4 VRAM
~66.1 GB
Image saved from public model sources
Mistral Medium 3.5 128B on your hardware
Dense flagship Mistral model for instruction following, reasoning, coding agents, and multimodal work. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.
128B total parameters. Dense models use the whole network for each token.
Start with Q4_K_M: about 62.6 GB on disk and 66.1 GB VRAM before extra context and runtime overhead.
chat, vision, reasoning, code
Listed as Modified MIT from huggingface.co/mistralai/Mistral-Medium-3.5-128B.
Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.
Open HF repoBenchmark snapshot
Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.
Agentic
Tau3-Telecom
91.4%
Coding agent
SWE-Bench Verified
77.6%