Meituan has released LongCat-Flash-Lite-Sparse, a new model featuring an Mixture of Experts (MoE) architecture with approximately 3 billion active parameters. This model utilizes a 30 billion n-gram lookup table that is offloaded to RAM, enabling a fast 256,000 token context window on a 24GB GPU. The approach is compared to Gemma 4's PLE trick, though initial analysis suggests it may not outperform existing models like Qwen 3.6 27b. AI
IMPACT Offers a large context window on consumer hardware, potentially enabling new applications for local LLMs.
RANK_REASON Release of a new model from a company that is not a tier-1 frontier lab. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →