back

GLM-5.3

GLM-5.3

Z.ai · 753B (40B active) · Mixture of Experts

Z.ai flagship open-weight coding model built for long-horizon agents and 1M-token project context

0 downloads 0 likes 2026-08 1024K context

Use Cases

chat reasoning code

Quantization Options

Quant Bits VRAM Quality Status
Q2_K 2 241.6 GB low
Q3_K_M 3 338 GB moderate
Q4_K_M 4 386.2 GB good
Q5_K_M 5 482.6 GB good
Q6_K 6 579.1 GB excellent
Q8_0 8 771.9 GB excellent
F16 16 1543.3 GB lossless

About this model

HF model card GLM
GLM-5.3 source artwork

Z.ai

753B

Mixture of Experts model for chat, reasoning, code workloads.

Context

1024K

Q4 VRAM

~386.2 GB

Image saved from public model sources

Mixture of Experts Thinking Tool use Code

GLM-5.3 on your hardware

Z.ai flagship open-weight coding model built for long-horizon agents and 1M-token project context. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.

Model shape

753B total parameters with 40B active. Dense models use the whole network for each token.

Local fit target

Start with Q4_K_M: about 368.2 GB on disk and 386.2 GB VRAM before extra context and runtime overhead.

Best use cases

chat, reasoning, code

License and source

Listed as GLM-5.3 from huggingface.co/zai-org/GLM-5.3.

Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.

Open HF repo