Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

Chronological Source Flow
Back

AI Fusion Summary

Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. This 320B-total and 18B-active MoE features a 1,048,576-token context window and MIT-licensed weights on Hugging Face. It utilizes hybrid KDA linear and NoPE sparse MLA attention to reduce compute and KV cache. LM Studio has integrated GLM-5.3-Flash into its Bionic AI agent platform, offering multimodal input and lower pricing compared to GLM-5.2, with API costs at $0.15/M input and $0.50/M output.
Community Comments
Loading updates...
0