A new research paper published on arXiv explores occupational bias within language models, revealing that these models often retain internal associations with demographic attributes like gender, race, and socioeconomic status, even when behavioral metrics do not detect disparities. The study introduces a causal framework to measure these biases, distinguishing between a model's internal representation of user competence and its observable outputs. Researchers demonstrated that by intervening on these internal representations, they could causally influence model behavior in tasks such as question-answering and hiring, suggesting potential failure modes that standard behavioral evaluations might miss. AI
IMPACT Reveals that language models may harbor hidden biases affecting their internal representations of user competence, even when behavioral outputs appear unbiased.
RANK_REASON The cluster contains a research paper detailing a mechanistic analysis of occupational bias in language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Keren De Jesus Fuentes
- Language Models
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →