PulseAugur
中
实时 21:36:22
English(EN) cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

新的 cMoLLM 架构通过动态卷积缩放 LLM

研究人员推出了一种新颖的大型语言模型缩放方法 cMoLLM,该方法将混合专家(MoE)风格融入整个模型管道,而不仅仅是前馈网络。该方法将 MoE 层重新构建为动态卷积,允许进行条件于输入的核聚合以及端到端流的可微分路由。在 FineWeb 数据集上使用 GPT-2 风格模型进行的实验表明,在匹配的计算预算下,cMoLLM 提高了语言建模困惑度和下游任务的准确性,与 ParaScale 和 AltUp 等现有方法相比,显示出更稳定的优化和更好的流利用率。 AI

影响 引入了一种新颖的 LLM 缩放策略,可能导致更大模型的更高效训练和推理。

排序理由 该集群包含一篇详细介绍 LLM 新模型架构和缩放定律的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 cMoLLM 架构通过动态卷积缩放 LLM

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍 LLM 新模型架构和缩放定律的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
72 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Xin Yang, Yemin Wang, Mingda Liu, Letian Li, Shuaishuai Cao, Zhengxiao He, Ryan Dong ·

    cMoLLM规模化:面向Mixture-of-LLMs的水平扩展定律

    arXiv:2607.22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a…