Researchers have developed IslamicTurathBench (ISTB), a new dataset designed to evaluate the capabilities of large language models (LLMs) in understanding classical Islamic scholarship. ISTB comprises 3,465 question-answer items derived from 35 recognized scholarly works spanning over 12 centuries and covering seven key fields within Islamic Studies. The dataset is structured to assess LLMs across different levels of scholarly demand and various task formats, including multiple-choice questions, passage comprehension, and open-ended knowledge questions, aiming to provide a comprehensive profiling of model performance in this specialized domain. AI
IMPACT This benchmark could lead to more accurate and culturally sensitive LLM performance in specialized academic and religious domains.
RANK_REASON The cluster describes a new academic benchmark dataset for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →