PulseAugur
中
实时 11:11:56

研究人员开发从维基媒体转储创建高质量训练语料库的方法

研究人员开发了一种方法,可以从原始维基媒体转储中为七种南斯拉夫语创建高质量的训练语料库。该过程包括两个主要阶段:从各种维基百科项目中提取和清理文本,然后使用基于n-gram的策略过滤掉低质量或重复的文章。这种方法旨在生成适合训练语言模型和进行比较语言学研究的语言丰富的数据集,并有可能推广到其他语言。 AI

影响 提供了一种生成专业语言语料库的可扩展方法,有可能提高大型语言模型在资源匮乏语言上的性能。

排序理由 详细介绍创建训练数据方法的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

研究人员开发从维基媒体转储创建高质量训练语料库的方法

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
详细介绍创建训练数据方法的学术论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
163 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Mihailo \v{S}kori\'c ·

    Wiki Dumps to Training Corpora: South Slavic Case

    arXiv:2604.25384v1 Announce Type: new Abstract: This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from ra…

  2. arXiv cs.CL TIER_1 English(EN) · Mihailo Škorić ·

    Wiki Dumps to Training Corpora: South Slavic Case

    This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw dumps of Wikipedia, Wikisource, Wikibooks, Wik…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    Wiki Dumps to Training Corpora: South Slavic Case

    This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw dumps of Wikipedia, Wikisource, Wikibooks, Wik…