A new demonstration, Maple-Preview, showcases a 20-billion parameter Mixture-of-Experts (MoE) model capable of running at 120 tokens per second on an iPhone. This development highlights the increasing efficiency and on-device capabilities of large language models. AI
IMPACT Demonstrates significant progress in on-device AI processing, potentially enabling more powerful mobile applications.
RANK_REASON Demonstration of a model running on consumer hardware, not a new model release from a frontier lab.
Read on Hacker News — AI stories ≥50 points →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →