Ling 3.0 Flash, a 124B parameter Mixture of Experts model from Ant Group's inclusionAI, has been released with MIT license and is available on Hugging Face. This model is designed for local execution, requiring significantly less memory than other large models like Kimi K3 due to its architecture, which activates only 5.1B parameters per token. It supports long contexts with hybrid linear attention and can perform reasoning by default, with the option to disable it for faster responses. The model can run interactively on Macs with 96GB of unified memory or on systems with a 24GB GPU and substantial system RAM, with community GGUF conversions available. AI
IMPACT Enables local execution of large models, potentially lowering barriers for researchers and developers.
RANK_REASON Model release from a major AI lab (Ant Group's inclusionAI) with detailed technical specifications and availability. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Ant Group
- DeepInfra
- Hugging Face
- inclusionAI
- Kimi k3
- Ling-3.0-flash
- llama.cpp
- Locally Uncensored
- Mac
- OpenRouter
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →