PulseAugur
EN
LIVE 07:35:33

ML Peer Review Scores Incomparable Across Topics, Study Finds

A new position paper published on arXiv argues that reviewer scores in Machine Learning peer review are not comparable across different research areas. The paper analyzes data from the International Conference on Learning Representations (ICLR) from 2021-2026, revealing that a paper's acceptance probability can vary by up to eight times based on its topic, even with the same reviewer score. The authors attribute this discrepancy to a measurement design failure rather than individual bias, suggesting that fixed numerical scales aggregate quality judgments across communities with non-uniform reviewer pools. They propose adopting calibrated review signals and publishing topic-stratified acceptance rates as a fairness metric. AI

IMPACT Highlights a potential systemic bias in ML conference peer review, suggesting a need for improved evaluation metrics.

RANK_REASON This is a research paper published on arXiv discussing a methodology issue in ML peer review. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

ML Peer Review Scores Incomparable Across Topics, Study Finds

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Binyan Xu, Fan Yang, Xilin Dai, Kehuan Zhang ·

    Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

    arXiv:2607.27209v1 Announce Type: cross Abstract: Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instru…