A new study published on arXiv investigates the impact of politeness in prompts on large language models used as relevance judges. Researchers found that tone significantly affects model judgments, with effects varying by model. When tone influences agreement, it appears to shift the model's overall scoring leniency rather than improving judgment accuracy. The study suggests that prompt tone can be a validity threat when precise relevance labels are critical. AI
IMPACT Highlights a potential validity threat in LLM-based evaluation metrics, impacting the reliability of AI-generated relevance judgments.
RANK_REASON Academic paper on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →