Zhipu AI has released GLM-5.3-Flash, a 320 billion parameter model with 18 billion active parameters and native image-text modality. This model utilizes a hybrid architecture combining sparse and linear attention mechanisms for efficient long-context serving. It supports FP8 and is available under the MIT license, deployable via SGLang and vLLM. AI
IMPACT Introduces a new frontier model with native multimodal capabilities and efficient long-context serving, potentially impacting multimodal AI development.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
- linear attention
- Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation
- Zhipu AI
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →