A new benchmark for narrative infilling, designed to evaluate how well Large Language Models can reconstruct missing sentences in stories, has been introduced. The benchmark, comprising approximately 9.2K instances across four narrative types, was used to test 20 open-source LLMs. Surprisingly, model scale did not correlate with performance; Gemma-2-2B achieved the highest qualitative score, surpassing larger models like DeepSeek-Qwen-32B and LLaMA-3.3-70B. Explicit reasoning techniques provided only marginal improvements, suggesting that narrative characteristics and length are more significant factors in task difficulty for current LLMs. AI
IMPACT This research highlights that smaller, more efficient models can achieve superior performance on complex tasks, potentially influencing future LLM development and deployment strategies.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- DeepSeek-Qwen-32B
- Gemma 2-2B
- Gotit.pub
- Hugging Face
- LLaMA 3.3 70B Instruct
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →