Researchers have developed a new method called Quantile-Guided Density Estimation (QGDE) to estimate the composition of hidden training corpora for large language models (LLMs). This technique leverages released tokenizer vocabularies, showing that token ID-to-ratio distributions are stable across different corpora. QGDE approximates these distributions using quantile trends and local density weighting, achieving low estimation errors in controlled and real-world scenarios, including with the SmolLM tokenizer. The findings suggest that tokenizer vocabularies can offer valuable insights into fine-grained corpus estimation, going beyond broad mixture inferences. AI
IMPACT Provides a novel method for analyzing LLM training data composition, potentially impacting model interpretability and bias detection.
RANK_REASON Academic paper detailing a new method for analyzing LLM training data. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →