Researchers have developed OrderProbe, a new benchmark designed to evaluate how well large language models (LLMs) can reconstruct the precise structural order of text. Unlike previous methods that allowed for multiple correct reorderings, OrderProbe uses fixed four-character expressions in Chinese, Japanese, and Korean to enable exact-match scoring. Experiments on twelve LLMs revealed that even advanced models struggle with this task, often achieving less than 35% accuracy in zero-shot recovery, indicating a gap between semantic understanding and precise structural reconstruction. AI
IMPACT Highlights a key limitation in current LLMs, suggesting a need for improved architectural designs focused on structural integrity.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →