A new research paper published on arXiv explores how large language models (LLMs) respond to demographic cues, finding that different cues for the same demographic group can lead to inconsistent conclusions about personalization and bias. The study analyzed over 14.8 million prompts in a U.S. context, revealing that model responses varied significantly depending on the specific cue used (e.g., names, race, gender). This suggests that LLM behavior is more sensitive to the linguistic signals within cues rather than stable demographic categories, advocating for multi-cue evaluations to better understand demographic variation in LLM outputs. AI
IMPACT Highlights the need for more nuanced evaluation of LLMs to understand how they respond to demographic signals, impacting AI safety and fairness research.
RANK_REASON Research paper published on arXiv detailing findings about LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Large Language Models
- Manuel Tonneau
- ScienceCast
- U.S.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →