PulseAugur
中
实时 08:50:37

新方法将英语LLM数据分类器适配用于多语言预训练

研究人员开发了一种方法,用于将基于英语的质量分类器适配于多语言大型语言模型(LLM)的高质量预训练数据选择。该方法通过在Transformer编码器嵌入之上训练一个小型多层感知机来实现,使用机器翻译文本和英语分类器的分数作为标签。跨不同模型规模的实验表明,这种多语言适配在不损害区域和文化知识的情况下,保持了下游LLM基准性能。 AI

影响 通过利用现有的英语数据质量分类器,实现了更有效和高效的多语言LLM预训练。

排序理由 学术论文,详细介绍了LLM预训练数据选择的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新方法将英语LLM数据分类器适配用于多语言预训练

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了LLM预训练数据选择的新方法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Vinko Sabol\v{c}ec, Bettina Messmer, Yassine Turki, Martin Jaggi ·

    为多语言大模型预训练数据选择调整英文质量分类器

    arXiv:2610.11585v1 Announce Type: cross Abstract: Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale…