The Ling team has released Ling 3.0 Tiny, a smaller version of their previous Ling 3.0 Flash model. This new model features 8 billion parameters with 1.3 billion active parameters, positioning its performance between Qwen and Gemma models in its size class. It is designed for high speed, achieving approximately 100-105 tokens/sec on DGX Spark and 86-90 tokens/sec on an M4 Pro MacBook with an 8K context length, while using around 8.34 GiB of memory. AI
IMPACT Offers a high-speed, smaller model option for local LLM deployments, potentially improving efficiency for certain tasks.
RANK_REASON Release of a new, smaller MoE model with performance and speed metrics. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →