PulseAugur
中
实时 19:51:39

新数据集揭示大型语言模型难以理解印尼价值观

研究人员开发了一个名为 Pancasila-Dilemmas 的新数据集,用于评估大型语言模型(LLMs)在多大程度上符合印尼人的价值观。该数据集包含 1,834 个基于印尼新闻的问题,重点关注五个核心价值观:宗教、人道、统一、民主和社会公正。对 50 个大型语言模型的初步评估显示,所有模型的表现都很差,概率匹配得分(Probability Match Score)低于 0.5,最大投票一致性得分(Max-Vote Agreement Score)低于 0.72,尤其在处理与宗教和统一相关的困境时表现困难。 AI

影响 凸显了大型语言模型在非西方语境下的价值观对齐方面的差距,可能推动更具文化意识的人工智能发展。

排序理由 该集群包含一篇介绍大型语言模型新数据集和评估方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新数据集揭示大型语言模型难以理解印尼价值观

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇介绍大型语言模型新数据集和评估方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
79 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Supryadi, Irfan, Julianti, Darren Keanly Martin, Jayvin Fernando, Yuqi Ren, Deyi Xiong ·

    Pancasila-Dilemmas: 在基于Pancasila的印度尼西亚人类价值观困境上评估大型语言模型

    arXiv:2607.18066v1 Announce Type: new Abstract: The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment focus on Western or universal values, while assessm…