PulseAugur
实时 08:45:02
English(EN) CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

新的CrossLex基准测试LLM的跨司法管辖区法律推理能力

研究人员推出了CrossLex,这是一个旨在评估大型语言模型在不同司法管辖区内法律推理能力的新基准。该基准使用相同的案件事实,但应用了中国、加利福尼亚州和德国特有的法律规则,涵盖了五个法律领域。CrossLex包含超过6000个带有法律答案和引用的实例,并提出了一个新的名为“Grounded Joint”的指标来评估正确性和来源约束性。初步实验表明,虽然当前的LLM通常可以在单一司法管辖区内准确回答法律问题,但它们在跨司法管辖区比较方面存在困难。 AI

影响 该基准有望推动LLM处理复杂、特定司法管辖区法律任务的能力的提升。

排序理由 该集群描述了一篇介绍用于评估LLM能力的新基准的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的CrossLex基准测试LLM的跨司法管辖区法律推理能力

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Xiaocui Yang, Xican Tan, Shoujie Chen, Shihan Xiao, Keke Tong, Xinyu Zhou ·

    CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

    arXiv:2608.01292v1 Announce Type: new Abstract: Legal reasoning is inherently jurisdiction-dependent: the same facts can call for different legal rules and yield different conclusions across legal systems. Yet existing benchmarks rarely evaluate whether large language models (LLM…