A new research paper published on arXiv explores the phenomenon of relative selection bias in language models, particularly how human corrections can amplify this bias. The study analyzes localized gradients to understand the interaction between edited and retained components of text, finding that localization can increase or decrease bias depending on specific factors. While importance weighting can help recover population means, the paper identifies issues with differing targets and derives metrics for mean-squared error and optimal retention coefficients. The research uses public human-post-edit experiments and synthetic selection to demonstrate these effects, distinguishing between relative and absolute bias. AI
IMPACT Provides a deeper understanding of bias amplification in language models, potentially informing future methods for data curation and model training.
RANK_REASON The cluster contains a single academic paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →