A new paper highlights a significant discrepancy in how large language model training datasets are measured, specifically concerning web-PDF corpora. Researchers found that statistics often report corpus size in tokens but calculate metrics per document, leading to a "unit bias." This bias means that a small fraction of documents, particularly those over 50 pages or generated by TeX toolchains, contain a disproportionately large amount of text. The study also reveals that common truncation caps, like the 5 MiB limit, result in substantial data loss, with widely used libraries failing to recover a significant portion of the truncated text. The authors recommend reporting corpus statistics in both document and token units to provide a more accurate representation of dataset composition. AI
IMPACT Highlights critical data quality issues in LLM training datasets, potentially impacting model performance and bias.
RANK_REASON The cluster contains an academic paper detailing a new finding about data statistics. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →