A new engine called Kyojin, built on ExLlamaV3 and optimized for AMD's Strix Halo hardware, has been developed to run large Mixture-of-Experts (MoE) models. This engine allows two 300B-class MoE models, GLM-5.3-Flash and MiMo-V2.6-Flash, to each fit within a single 128 GB machine. Performance benchmarks show GLM-5.3-Flash achieving approximately 580 tokens/s for prefill and MiMo-V2.6-Flash reaching up to 44 tokens/s for decoding, with competitive quality metrics compared to official FP8 versions. AI
IMPACT Enables running large MoE models on more accessible hardware, potentially democratizing advanced AI capabilities.
RANK_REASON Release of a new engine and quantized models for local LLM deployment on specific hardware. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →