DeepSeek · 1.65T (49B active) · Mixture of Experts
Frontier DeepSeek V4 MoE for coding, reasoning, and million-token long-context workloads
Use Cases
Mixture of Experts
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 528.7 GB | low | — |
| Q3_K_M | 3 | 740 GB | moderate | — |
| Q4_K_M | 4 | 845.7 GB | good | — |
| Q5_K_M | 5 | 1057 GB | good | — |
| Q6_K | 6 | 1268.3 GB | excellent | — |
| Q8_0 | 8 | 1690.9 GB | excellent | — |
| F16 | 16 | 3381.2 GB | lossless | — |
About this model
DeepSeek
Mixture of Experts model for chat, reasoning, code workloads.
Context
1024K
Q4 VRAM
~845.7 GB
Image saved from public model sources
DeepSeek V4 Pro 0813 on your hardware
Frontier DeepSeek V4 MoE for coding, reasoning, and million-token long-context workloads. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.
1.65T total parameters with 49B active. 6 of 384 experts are active per token.
Start with Q4_K_M: about 806.8 GB on disk and 845.7 GB VRAM before extra context and runtime overhead.
chat, reasoning, code
Listed as MIT from huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813.
Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.
Open HF repoBenchmark snapshot
Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.
Coding agent
Terminal Bench 2.1
87.9
Coding agent
DeepSWE
62.7
Tool use
Toolathlon-Verified
74.1
Full-stack agent
DSBench-FullStack
71.1
Full-stack agent
DSBench-Hard
67.2
Agentic work
AutomationBench
31.8
Knowledge
HLE
42.7 / 60.0
Without / with tools
Codebase generation
NL2Repo
61.5