Moonshot AI · 2.8T (16/896 experts active) · Mixture of Experts
Native multimodal frontier MoE for long coding sessions, research, vision, and agentic knowledge work
Use Cases
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 896.9 GB | low | — |
| Q3_K_M | 3 | 1255.5 GB | moderate | — |
| Q4_K_M | 4 | 1434.7 GB | good | — |
| Q5_K_M | 5 | 1793.3 GB | good | — |
| Q6_K | 6 | 2151.9 GB | excellent | — |
| Q8_0 | 8 | 2869 GB | excellent | — |
| F16 | 16 | 5737.4 GB | lossless | — |
About this model
Moonshot AI
Mixture of Experts model for chat, vision, reasoning workloads.
Context
1024K
Q4 VRAM
~1434.7 GB
Image saved from public model sources
Kimi K3 on your hardware
Native multimodal frontier MoE for long coding sessions, research, vision, and agentic knowledge work. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.
2.8T total parameters with 16/896 experts active. Dense models use the whole network for each token.
Start with Q4_K_M: about 1369 GB on disk and 1434.7 GB VRAM before extra context and runtime overhead.
chat, vision, reasoning, code
Listed as Kimi from huggingface.co/moonshotai/Kimi-K3.
Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.
Open HF repoBenchmark snapshot
Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.
Reasoning
GPQA Diamond
93.5
Coding
Terminal-Bench 2.1
88.3
Coding
DeepSWE
67.5
Agentic search
BrowseComp
91.2
Tool use
Toolathlon-Verified
76.5
MCP/tool use
MCPMark-Verified
94.5
Computer use
OSWorld-Verified
84.8
Long context
AA-LCR
74.7