The Terminal Bench v4 benchmark results show GLM-5.3 as the top-performing open model, significantly outperforming others in its class. GLM-5.3-Flash also leads among flash models, while Kimi-K3 performed poorly relative to its size. Qwen3.8-27B is noted as the only small model capable of achieving a score on this benchmark. AI
IMPACT Provides a comparative performance analysis of various open-source LLMs, highlighting GLM-5.3's leading capabilities.
RANK_REASON Benchmark results for open-source LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- DSV4.1-Flash
- DSV4-Flash
- DSV4-Pro
- gemma4-31b
- GLM-5.3
- Kimi-K3
- Muse Glimmer
- Qwen3.8-27B
- Qwen3.8-Flash-Next
- Terminal Bench v4
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →