A new position paper published on arXiv argues that reviewer scores in Machine Learning peer review are not comparable across different research areas. The paper analyzes data from the International Conference on Learning Representations (ICLR) from 2021-2026, revealing that a paper's acceptance probability can vary by up to eight times based on its topic, even with the same reviewer score. The authors attribute this discrepancy to a measurement design failure rather than individual bias, suggesting that fixed numerical scales aggregate quality judgments across communities with non-uniform reviewer pools. They propose adopting calibrated review signals and publishing topic-stratified acceptance rates as a fairness metric. AI
IMPACT Highlights a potential systemic bias in ML conference peer review, suggesting a need for improved evaluation metrics.
RANK_REASON This is a research paper published on arXiv discussing a methodology issue in ML peer review. [lever_c_demoted from research: ic=1 ai=1.0]
- area chairs
- arXiv
- community priors
- expert reviewer standards
- ICLR 2021--2026
- International Conference on Learning Representations
- ML Peer Review
- program committees
- quality dilution
- reviewer scores
- review signals
- scoring culture
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →