Researchers have introduced CrossLex, a new benchmark designed to evaluate how well large language models can perform legal reasoning across different jurisdictions. The benchmark uses identical fact patterns but applies legal rules specific to China, California, and Germany, covering five areas of law. CrossLex includes over 6,000 instances with legal answers and citations, and proposes a new metric called Grounded Joint to assess both correctness and source grounding. Initial experiments indicate that while current LLMs can often answer legal questions accurately within a single jurisdiction, they struggle with cross-jurisdictional comparisons. AI
IMPACT This benchmark could drive improvements in LLMs' ability to handle complex, jurisdiction-specific legal tasks.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →