Two new research papers explore how large language models (LLMs) encode and perpetuate stereotypes. The first, STEREODISCO, uses a framework adapted from social psychology to identify stereotypical axes in LLM internal representations, finding that models like LLaMA-3-8B-Instruct and Mistral-7B-Instruct agree with each other more than with human perceptions on social group stereotypes. The second paper introduces the Stereotypes-to-Decisions (S2D) framework to evaluate regional bias in LLMs, specifically focusing on China, and reveals that these models exhibit systematic regional biases in perceptions of warmth and competence, which correlate with regional development indicators and remain stable across different language prompts. AI
IMPACT Highlights the need for more nuanced evaluation of LLMs beyond performance metrics, focusing on their societal biases and potential for perpetuating harmful stereotypes.
RANK_REASON Two academic papers published on arXiv detailing new methodologies for evaluating stereotypes and regional bias in LLMs.
Read on Hugging Face Daily Papers →
- arXiv
- China
- competency
- Education
- English
- large language models (LLMs)
- Occupation
- social interaction
- Standard Chinese
- Stereotypes-to-Decisions (S2D)
- Warmth
- LLaMA-3-8B-Instruct
- LLMs
- Mistral-7B-Instruct
- Osgood
- social psychology
- STEREODISCO
- WordNet
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →