Researchers have developed a new method called Sampled-BPE to efficiently audit large Chinese web corpora for language model pollution. This technique significantly reduces runtime and memory usage compared to full scans, while maintaining accurate estimates of pollution categories. The pipeline was applied to numerous open Chinese corpora and Common Crawl snapshots, revealing widespread and temporally varying pollution. A detailed dataset of Chinese web tokens, including context and category information, has also been released to aid further analysis and tracing of pollution. AI
IMPACT Provides a more efficient way to clean and verify data used for training large language models.
RANK_REASON Academic paper detailing a new methodology for corpus auditing. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →