PulseAugur
EN
LIVE 04:57:45

Polite prompts alter LLM relevance judgments, study finds

A new study published on arXiv investigates the impact of politeness in prompts on large language models used as relevance judges. Researchers found that tone significantly affects model judgments, with effects varying by model. When tone influences agreement, it appears to shift the model's overall scoring leniency rather than improving judgment accuracy. The study suggests that prompt tone can be a validity threat when precise relevance labels are critical. AI

IMPACT Highlights a potential validity threat in LLM-based evaluation metrics, impacting the reliability of AI-generated relevance judgments.

RANK_REASON Academic paper on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Polite prompts alter LLM relevance judgments, study finds

How we ranked this

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Academic paper on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Meng Li ·

    Should I Be Polite to My LLM Relevance Judge? Tone as a Severity Operating-Point Shift

    Large language models are increasingly used as relevance judges, yet their labels can shift with prompt surface form. We study one such feature -- tone -- on 3,498 TREC DL19/DL20 query-passage pairs, across eight judge models, five classifier-calibrated politeness levels, and thr…