PulseAugur
中
实时 23:20:28
English(EN) Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models

专家混合(Mixture-of-Experts)架构主导大型语言模型发布,带来效率提升 · 已追踪2个来源

专家混合(Mixture-of-Experts, MoE)架构正日益主导大型语言模型领域,在2026年部署的七个最广泛使用的开源模型中,有六个采用了这种方法。该架构最早于2017年提出,它允许模型拥有海量参数(代表知识容量),但每次仅激活一小部分来处理输入令牌,从而显著降低了每令牌的计算成本。DeepSeek V4.1 Flash和Xiaomi MiMo-V2.6 Pro等近期发布的模型展示了这种效率,由于其稀疏激活,它们以更低的每令牌价格提供了前沿水平的智能。然而,其权衡是需要大量的内存,因为无论是否激活,所有专家都必须驻留在VRAM中。 AI

影响 专家混合(Mixture-of-Experts)模型的广泛采用表明,大型语言模型正朝着计算效率更高的方向发展,这可能降低推理成本并支持更大规模的模型。

排序理由 该集群讨论了近期大型语言模型发布背后的技术架构(专家混合,Mixture-of-Experts)及其影响,并得到了研究论文和模型细节的支持。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

专家混合(Mixture-of-Experts)架构主导大型语言模型发布,带来效率提升 · 已追踪2个来源

本文如何被排名

Signal score
84 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群讨论了近期大型语言模型发布背后的技术架构(专家混合,Mixture-of-Experts)及其影响,并得到了研究论文和模型细节的支持。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [2]

  1. dev.to — LLM tag TIER_1 English(EN) · Daniel Sam Pete Thiyagu ·

    Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models

    <p>Two open-weight releases, twelve days apart. <strong>DeepSeek V4.1 Flash</strong> (September 10): 552 billion parameters — and about 8 billion of them fire on each input token, 16 billion on output. <strong>Xiaomi MiMo-V2.6 Pro</strong> (September 22): 1.02 <em>trillion</em> p…

  2. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models # ai # machinelearning # deeplearning # llm # software # coding

    Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models # ai # machinelearning # deeplearning # llm # software # coding # development # engineering # inclusive # community Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behi…