Z.ai has released GLM-5.3-Flash, a new 320-billion-parameter mixture-of-experts model that is natively multimodal and open-source. This model features a hybrid attention architecture combining sparse and linear attention, enabling it to handle a 1-million-token context window efficiently. Z.ai claims GLM-5.3-Flash outperforms its predecessor on coding and agentic benchmarks at a significantly lower inference cost, while also approaching the performance of Claude Opus 4.8. AI
IMPACT Sets new SOTA on coding benchmarks at significantly lower cost, potentially accelerating adoption of multimodal and long-context models.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →