PulseAugur
EN
LIVE 23:37:22

Meituan releases LongCat-Flash-Lite-Sparse with 256k context

Meituan has released LongCat-Flash-Lite-Sparse, a new model featuring an Mixture of Experts (MoE) architecture with approximately 3 billion active parameters. This model utilizes a 30 billion n-gram lookup table that is offloaded to RAM, enabling a fast 256,000 token context window on a 24GB GPU. The approach is compared to Gemma 4's PLE trick, though initial analysis suggests it may not outperform existing models like Qwen 3.6 27b. AI

IMPACT Offers a large context window on consumer hardware, potentially enabling new applications for local LLMs.

RANK_REASON Release of a new model from a company that is not a tier-1 frontier lab. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Meituan releases LongCat-Flash-Lite-Sparse with 256k context

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Gohab2001 ·

    Meituan just dropped LongCat-Flash-Lite-Sparse

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbsztw/meituan_just_dropped_longcatflashlitesparse/"> <img alt="Meituan just dropped LongCat-Flash-Lite-Sparse" src="https://external-preview.redd.it/TB8lvc_WvC59pV7HAeNN_QTT2Sa6SPRzddgrjEvNFGA.png?width=640&…