Researchers have developed a new pipeline for curating high-quality web corpora specifically for European Portuguese (PT-PT). This pipeline addresses challenges like dialectal overlap with Brazilian Portuguese and the scale of data processing. It efficiently processes 411 TB of raw data, incorporating a novel post-scraping block that improves document yield by 19.04% by preventing premature discarding of valid text. The framework includes rigorous language identification, weighted fuzzy deduplication, and neural quality classification, resulting in a clean and representative corpus suitable for LLM pre-training. AI
IMPACT Provides a scalable framework and a clean corpus optimized for LLM pre-training, potentially improving model performance on European Portuguese.
RANK_REASON The item is a research paper detailing a new method for data curation. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Brazilian Portuguese
- European Portuguese
- Fine PT-PT Web
- Gonçalo Vinagre
- GV.Martins
- Hugging Face
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →