PulseAugur
实时 09:59:36
English(EN) Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

新流程为LLM精心策划了包含410亿token的欧洲葡萄牙语网页语料库

研究人员开发了一个新的流程,用于精心策划专门针对欧洲葡萄牙语(PT-PT)的高质量网页语料库。该流程解决了巴西葡萄牙语的方言重叠和数据处理规模等挑战。它高效地处理了411 TB的原始数据,并包含了一个新颖的抓取后区块,通过防止过早丢弃有效文本,将文档产出率提高了19.04%。该框架包括严格的语言识别、加权模糊去重和神经质量分类,最终形成了一个干净且具有代表性的语料库,适用于LLM的预训练。 AI

影响 提供了一个可扩展的框架和一个为LLM预训练优化的干净语料库,有可能提高模型在欧洲葡萄牙语上的性能。

排序理由 该项目是一篇研究论文,详细介绍了一种新的数据策划方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新流程为LLM精心策划了包含410亿token的欧洲葡萄牙语网页语料库

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目是一篇研究论文,详细介绍了一种新的数据策划方法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Gon\c{c}alo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos, Duarte Miguel Alves, Afonso Simpl\'icio, Diogo Tavares, David Semedo, Daniel Gomes, Jo\~ao Magalh\~aes ·

    Fine PT-PT Web:高质量的 410 亿词元欧洲葡萄牙语网络数据集

    arXiv:2609.07699v2 Announce Type: cross Abstract: Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a…