Researchers have developed a new benchmark to assess large language models' (LLMs) ability to determine the correct version of a law based on when legal facts occurred. Their experiments revealed that LLMs tend to favor the most recent law, even when it's not applicable to the case's timeline. This bias appears to stem from reinforcement learning shaping explicit reasoning, which can limit the diversity of reasoning paths and cause models to over-rely on current statutes. Counterintuitively, models with stronger general reasoning capabilities performed worse on this specific temporal legal reasoning task. AI
IMPACT Highlights a specific failure mode in LLMs for temporal reasoning, suggesting improvements are needed for applications requiring historical legal context.
RANK_REASON Academic paper detailing a new benchmark and findings on LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →