Researchers have developed SDUs DAISY, a new benchmark designed to evaluate large language models' understanding of Danish cultural heritage. The benchmark, derived from the Danish Culture Canon 2006 and Wikipedia, features 741 manually verified question-answer pairs covering a wide range of cultural artifacts. Initial testing revealed that current models, including Llama-3.3-70B and GPT-OSS-120B, performed poorly, indicating the benchmark's challenging nature. The dataset and evaluation tools are publicly available on Hugging Face Datasets. AI
IMPACT Highlights limitations in current LLMs' cultural knowledge and reasoning capabilities, potentially guiding future model development.
RANK_REASON The item is a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- 20b
- Daisy
- Danish Culture Canon 2006
- gemma3:27b
- GPT-OSS 120B
- Hugging Face Datasets
- Jacob Nielsen
- Llama 3.3 70B Instruct
- Mistral-3.1-24B
- schneiderkamplab
- SDUs DAISY
- Wikipedia
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →