PulseAugur
EN
LIVE 09:44:25

New benchmark dataset evaluates LLMs on Islamic scholarly tradition

Researchers have developed IslamicTurathBench (ISTB), a new dataset designed to evaluate the capabilities of large language models (LLMs) in understanding classical Islamic scholarship. ISTB comprises 3,465 question-answer items derived from 35 recognized scholarly works spanning over 12 centuries and covering seven key fields within Islamic Studies. The dataset is structured to assess LLMs across different levels of scholarly demand and various task formats, including multiple-choice questions, passage comprehension, and open-ended knowledge questions, aiming to provide a comprehensive profiling of model performance in this specialized domain. AI

IMPACT This benchmark could lead to more accurate and culturally sensitive LLM performance in specialized academic and religious domains.

RANK_REASON The cluster describes a new academic benchmark dataset for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark dataset evaluates LLMs on Islamic scholarly tradition

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Shahd Gaben, Heba Sbahi, Samer Rashwani, Abdessalam Bouchekif, Mutaz Al-Khatib, Emad Mohamed, Somaya Eltanbouly, Mohammed Ghaly ·

    IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

    arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key conce…