PulseAugur
EN
LIVE 23:36:59

LLMs as IR assessors show structured sensitivity to persona conditioning

A new research paper explores the use of persona conditioning to assess the sensitivity of large language models (LLMs) when used as relevance assessors in information retrieval (IR) evaluation. By instantiating five distinct assessor roles across six LLM backbones and two datasets (TREC DL20 and RAG24), the study found that LLM judgments exhibit structured sensitivity rather than uniform changes. While high-capacity models maintained system-ranking agreement, smaller models amplified persona-induced instability, with sensitivity concentrated on specific retrieval systems. AI

IMPACT This research provides a method to stress-test LLM evaluation pipelines, identifying systems sensitive to assessor framing and improving the reliability of IR evaluation.

RANK_REASON The cluster contains a research paper detailing a novel methodology for evaluating LLM performance in a specific domain (information retrieval).

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

LLMs as IR assessors show structured sensitivity to persona conditioning

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster contains a research paper detailing a novel methodology for evaluating LLM performance in a specific domain (information retrieval).
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
25 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini ·

    Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

    arXiv:2608.10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We stu…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Gianluca Demartini ·

    Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

    Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism …