zai-org/GLM-5.3-Flash

GLM-5.3-Flash, a 320B multimodal model with only 18B active parameters, is now publicly downloadable and supports deployment via SGLang, vLLM, Transformers and other runtimes. The release notes detail a hybrid sparse‑linear attention architecture and new scaling tricks, while providing API endpoints on Z.ai. Developers can run the model locally with standard libraries.

README

zai-org/GLM-5.3-Flash View on Hugging Face

Loading the README from Hugging Face…