inclusionAI has released Ling-3.0-tiny, a new hybrid reasoning Mixture-of-Experts (MoE) model with 7.9 billion total parameters and 1.3 billion activated parameters per token. This model is designed for efficient local deployment and resource-constrained environments, offering strong reasoning and agentic capabilities. It features a novel Kimi Delta Attention (KDA) and Multi-Head Latent Attention (MLA) architecture, supporting both fast responses and multi-step reasoning, and has been validated on hardware like Apple Silicon MacBooks. AI
IMPACT Enables more accessible local and edge AI deployments with advanced reasoning capabilities.
RANK_REASON Model release from a recognized AI lab (inclusionAI) with detailed technical specifications and performance metrics. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- Gemma
- Hugging Face
- inclusionAI
- Ling 3.0 Tiny
- Ling team
- Qwen
- Apple M4 Pro
- Apple Silicon MacBook
- inclusionAI/Ling-3.0-tiny
- Ling-3.0
- Mac mini
- Mixture-of-Experts
- modelscope
- NVIDIA DGX Spark
- OpenRouter
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →