A new research paper explores the use of persona conditioning to assess the sensitivity of large language models (LLMs) when used as relevance assessors in information retrieval (IR) evaluation. By instantiating five distinct assessor roles across six LLM backbones and two datasets (TREC DL20 and RAG24), the study found that LLM judgments exhibit structured sensitivity rather than uniform changes. While high-capacity models maintained system-ranking agreement, smaller models amplified persona-induced instability, with sensitivity concentrated on specific retrieval systems. AI
IMPACT This research provides a method to stress-test LLM evaluation pipelines, identifying systems sensitive to assessor framing and improving the reliability of IR evaluation.
RANK_REASON The cluster contains a research paper detailing a novel methodology for evaluating LLM performance in a specific domain (information retrieval).
Read on arXiv cs.IR (Information Retrieval) →
- DL 208
- NVIDIA Nemotron-Personas-USA
- PersonaHub
- RAG24
- TREC DL20
- umbrella
- alphaXiv
- CatalyzeX Code Finder for Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- ScienceCast
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →