A new evaluation framework called DiSCo has been developed to measure cultural preference bias in large language models (LLMs). Unlike previous benchmarks, DiSCo uses a distribution-first approach with forced-choice questions to isolate default cultural priors and test steerability. Evaluations using DiSCo-Bench, which includes items from 12 cultures, revealed that LLMs heavily favor UK and US cultural preferences, with prompt-based steering exacerbating this bias rather than resolving it. AI
IMPACT Highlights the need for more culturally equitable LLM development and evaluation, potentially influencing future model training and fine-tuning.
RANK_REASON The cluster contains an academic paper detailing a new evaluation framework for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →