Researchers have developed a new dataset called Pancasila-Dilemmas to evaluate how well large language models (LLMs) align with Indonesian human values. The dataset contains 1,834 questions based on Indonesian news, focusing on five core values: Religion, Humanity, Unity, Democracy, and Social Justice. Initial evaluations of 50 LLMs showed that all models performed poorly, achieving less than a 0.5 Probability Match Score and a 0.72 Max-Vote Agreement Score, with particular struggles in addressing dilemmas related to religion and unity. AI
IMPACT Highlights a gap in LLM value alignment for non-Western contexts, potentially driving development of more culturally-aware AI.
RANK_REASON The cluster contains a research paper introducing a new dataset and evaluation methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- Democracy
- Humanity
- Indonesian human values
- large language models
- Pancasila-Dilemmas
- Religion
- Social Justice
- Supryadi Supryadi
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →