PulseAugur
实时 22:33:43
English(EN) LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios

LLM 在新研究中显示出语言和人口统计学偏见

新研究表明,多语言大型语言模型在面对冲突信息时会表现出显著的语言偏见,通常偏袒某些语言而非其他语言。研究还揭示了 LLM 中性别、种族和年龄代表性的差异,而去偏见努力有时会产生新的公平性权衡。在职业和犯罪场景中,这些模型经常偏离现实世界的人口统计数据,并且它们的刻板印象会在不同语言中被放大。 AI

影响 强调了 LLM 中可能影响现实世界应用公平性和可靠性的关键偏见,需要改进缓解策略。

排序理由 多篇学术论文发表在 arXiv 上,详细介绍了 LLM 偏见评估。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

LLM 在新研究中显示出语言和人口统计学偏见

报道来源 [4]

  1. arXiv cs.CL TIER_1 English(EN) · Robert \"Ostling, Murathan Kurfal{\i} ·

    多语言大模型在冲突信息下的语言偏见

    arXiv:2604.07123v2 Announce Type: replace Abstract: Large Language Models (LLMs) have been shown to contain biases in the process of integrating conflicting information when answering questions. Here we ask whether such biases also exist with respect to which language is used for…

  2. arXiv cs.AI TIER_1 English(EN) · Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto ·

    IndoBias:用于印尼语 LLM 偏见评估的双轨道文化基础基准

    arXiv:2606.01260v1 Announce Type: cross Abstract: Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and locali…

  3. arXiv cs.AI TIER_1 English(EN) · Vishal Mirza, Rahul Kulkarni, Aakanksha Jadhav ·

    LLM偏见评估:职业和犯罪场景中的性别、种族和年龄差异

    arXiv:2409.14583v4 Announce Type: replace Abstract: LLM bias evaluation is critical as large language models (LLMs) increasingly influence high-stakes decisions. This paper provides a comprehensive assessment of gender, racial, and age disparities in leading LLMs, revealing that …

  4. arXiv cs.CL TIER_1 English(EN) · Jiwoo Choi, Seonwoo Ahn, Tongxin Zhang, Seohyon Jung ·

    将大语言模型性别偏见锚定至人类基线:一项跨语言审计

    arXiv:2605.30804v1 Announce Type: new Abstract: We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for English-language use (Claude, GPT, Gemini) and three for East Asian use (DeepSeek, S…