PulseAugur
实时 12:02:26
English(EN) A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models

MoE语言模型在双语训练中展现语言专业化

研究人员探讨了专家混合(MoE)语言模型如何为双语数据开发专业化路由。他们使用声明式-过程式框架,分析了一个在顺序语言暴露下训练的英语-德语MoE Transformer。研究发现,虽然经过课程训练的模型表现出一定的语言专业化,但在混合数据上训练的基线模型表现出更强的总体专业化,尽管这种专业化依赖于种子且集中在单一语言上。然而,课程方法产生了更稳定、语言平衡的路由配置文件。 AI

影响 为理解MoE模型的内部工作机制提供了见解,可能指导未来多语言能力的架构改进。

排序理由 学术论文,详细介绍了一种分析MoE模型行为的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

MoE语言模型在双语训练中展现语言专业化

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Amrit Gopinath (Sri Sivasubramaniya Nadar College of Engineering, Chennai, India), Raghul (Sri Sivasubramaniya Nadar College of Engineering, Chennai, India), Durairaj Thenmozhi (Shiv Nadar University Chennai, India) ·

    面向双语专家混合语言模型的专家路由的声明式-过程式视角

    arXiv:2608.15102v1 Announce Type: new Abstract: We investigate whether Mixture-of-Experts (MoE) language models develop linguistically structured expert routing during bilingual language acquisition. Inspired by the Declarative-Procedural framework, we analyze lexical, grammatica…