PulseAugur
实时 12:24:48
English(EN) LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

新基准评估人工智能生成的法律报告中的引文可信度

研究人员开发了LegalCiteTrust,一个旨在评估中文长篇法律研究报告中引文可信度的新基准。该基准从覆盖度、支持度和引文可信度三个维度评估报告,其中引文可信度又细分为存在性、保真度和适用性(E/F/A)。使用各种LLM和研究系统进行的实验表明,检索工具可以增强证据支持,但不能可靠地提高引文可信度。研究结果表明,可靠的法律研究生成需要引文感知治理,确保检索到的法律权威不仅被找到,而且被准确描述并恰当应用。 AI

影响 该基准有望推动用于法律研究的人工智能系统的可靠性和可信度方面的改进。

排序理由 该项目描述了一个用于评估特定领域人工智能系统的新学术基准。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准评估人工智能生成的法律报告中的引文可信度

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一个用于评估特定领域人工智能系统的新学术基准。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
60 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yunhan Li, Mingjie Xie, Zeyang Shi, Gengshen Wu, Min Yang ·

    LegalCiteTrust:中文长篇法律研究报告中的引文可信度基准测试

    arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, but also on whether cited legal authorities are trustworthy. A citation can be ri…