A new paper published on arXiv critiques the standard protocol for human evaluation of natural language generation (NLG) systems. The authors argue that common practices, particularly the use of Likert scales, can lead to inaccurate assessments of human preferences and even reverse the true direction of preference. They propose an alternative method called system-level probabilistic assessment (SPA) for evaluating open-ended tasks like story generation, demonstrating its effectiveness in correctly ordering GPT-3 models by size. AI
IMPACT Proposes a new evaluation protocol that could lead to more accurate assessments of NLG model capabilities.
RANK_REASON Research paper published on arXiv detailing a new evaluation protocol for NLG models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →