Alibaba · 2.4T (95B active) · Mixture of Experts
Qwen-Max-class open-weight MoE for complex coding, research, and long-horizon agent tasks
Use Cases
Mixture of Experts
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 768.8 GB | low | — |
| Q3_K_M | 3 | 1076.2 GB | moderate | — |
| Q4_K_M | 4 | 1229.8 GB | good | — |
| Q5_K_M | 5 | 1537.2 GB | good | — |
| Q6_K | 6 | 1844.5 GB | excellent | — |
| Q8_0 | 8 | 2459.2 GB | excellent | — |
| F16 | 16 | 4917.9 GB | lossless | — |
About this model
Alibaba
Mixture of Experts model for chat, reasoning, code workloads.
Context
256K
Q4 VRAM
~1229.8 GB
Image saved from public model sources
Qwen 3.8 2.4T-A95B on your hardware
Qwen-Max-class open-weight MoE for complex coding, research, and long-horizon agent tasks. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.
2.4T total parameters with 95B active. 11 of 512 experts are active per token.
Start with Q4_K_M: about 1173.5 GB on disk and 1229.8 GB VRAM before extra context and runtime overhead.
chat, reasoning, code
Listed as Apache 2.0 from huggingface.co/Qwen/Qwen3.8-2.4T-A95B.
Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.
Open HF repoBenchmark snapshot
Public eval numbers from the model card or benchmark indexes. Scores use each benchmark's own scale, so compare rows by task type, not as one combined rating.
Coding agent
Terminal Bench 2.1
86.6
Coding agent
SWE-bench Pro
67.7
Coding agent
DeepSWE 1.1
56.6
Research agent
PaperBench
93.0
Coding agent
QwenSWEBench
80.7
Tool use
Toolathlon Verified
72.5
Reasoning
GPQA Diamond
92.6
Long context
MRCR v2 256K
92.9
8-needle