Apple has developed a new architecture for its third-generation foundation models, dubbed AFM3 20B, which utilizes a technique called Instruction-Following Pruning. This method allows the model to activate approximately 20% of its MLP layers per prompt, significantly reducing the computational load. Unlike traditional models that might switch experts per token or layer, AFM3 20B routes prompts to specific experts, keeping a core set of shared experts always accessible while loading a smaller subset into memory. AI
IMPACT This architecture could lead to more efficient on-device AI processing by reducing active parameter usage.
RANK_REASON The cluster describes a novel model architecture and pruning technique, aligning with research into efficient AI model design. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →