PulseAugur
EN
LIVE 19:58:06

Ling 3.0 Tiny: New 8B MoE Model Offers High Speed

The Ling team has released Ling 3.0 Tiny, a smaller version of their previous Ling 3.0 Flash model. This new model features 8 billion parameters with 1.3 billion active parameters, positioning its performance between Qwen and Gemma models in its size class. It is designed for high speed, achieving approximately 100-105 tokens/sec on DGX Spark and 86-90 tokens/sec on an M4 Pro MacBook with an 8K context length, while using around 8.34 GiB of memory. AI

IMPACT Offers a high-speed, smaller model option for local LLM deployments, potentially improving efficiency for certain tasks.

RANK_REASON Release of a new, smaller MoE model with performance and speed metrics. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Ling 3.0 Tiny: New 8B MoE Model Offers High Speed

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/-Cubie- ·

    inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vkqwso/inclusionailing30tiny_8b_a13b_moe_hugging_face/"> <img alt="inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face" src="https://external-preview.redd.it/-W9E8wIes37LVjlOsySQwQoPCsCCBxgfutu0uAUT5WA.png…