back

GLM-5.3 Flash

MIT

Z.ai · 320B (18B active) · Mixture of Experts

Native multimodal GLM-5 Flash model with FP8 weights and a 1M-token context window

0 downloads 0 likes 2026-08 1024K context

Use Cases

chat vision reasoning code

Mixture of Experts

Total experts: 288
Active experts: 8
Active params: 18.0B

Quantization Options

Quant Bits VRAM Quality Status
Q2_K 2 102.9 GB low
Q3_K_M 3 143.9 GB moderate
Q4_K_M 4 164.4 GB good
Q5_K_M 5 205.4 GB good
Q6_K 6 246.4 GB excellent
Q8_0 8 328.3 GB excellent
F16 16 656.2 GB lossless

About this model

HF model card GLM
GLM-5.3 Flash source artwork

Z.ai

320B

Mixture of Experts model for chat, vision, reasoning workloads.

Context

1024K

Q4 VRAM

~164.4 GB

Image saved from public model sources

Mixture of Experts Thinking Tool use Vision Code

GLM-5.3 Flash on your hardware

Native multimodal GLM-5 Flash model with FP8 weights and a 1M-token context window. This page turns the Hugging Face model card into practical local-run numbers, so you can compare quantized VRAM, system RAM, and expected fit before downloading a large checkpoint.

Model shape

320B total parameters with 18B active. 8 of 288 experts are active per token.

Local fit target

Start with Q4_K_M: about 156.5 GB on disk and 164.4 GB VRAM before extra context and runtime overhead.

Best use cases

chat, vision, reasoning, code

License and source

Listed as MIT from huggingface.co/zai-org/GLM-5.3-Flash.

Want the real verdict? Pick your GPU or edit the specs on this page and compare the quant table below.

Open HF repo