Z.ai Releases GLM-5.3-Flash Under MIT: 320B Total, 18B Active, First Natively Multimodal GLM-5 Model
Z.ai published GLM-5.3-Flash on Hugging Face on August 26, 2026. The model runs 320B total parameters with 18B active, and the company says it beats GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks. The license is MIT.
Z.ai (zai-org) published GLM-5.3-Flash on Hugging Face on August 26, 20261. It is the first natively multimodal model in the GLM-5 series, configured with 320B total parameters and 18B active1. The company states that it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, and that it approaches Claude Opus 4.8 on coding and agentic benchmarks1. Both are Z.ai’s own claims rather than independently reproduced results.
The architecture was rebuilt starting from a newly trained base model, introducing for the first time in the GLM series a hybrid of sparse and linear attention, which the company says cuts long-context serving costs while keeping long-context accuracy intact1. It also adopts Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency, which together with a 30-trillion-token multimodal pre-training corpus is meant to deliver more capability per unit of compute1.
The license is MIT, with local deployment supported through SGLang, vLLM, TokenSpeed, and KTransformers alongside Z.ai’s own API1. The company also released GLM-5.3 on August 14, when it said it would hold the weights back until safety evaluation was complete.
Sources
- GLM-5.3-Flash - Z.ai (zai-org) official model card (August 26, 2026)
Was this article helpful?
Thank you!
Received. Thank you!