Researchers have developed LegalCiteTrust, a new benchmark designed to evaluate the trustworthiness of citations within Chinese long-form legal research reports. This benchmark assesses reports across three dimensions: Coverage, Support, and Citation Trustworthiness, with the latter further broken down into Existence, Fidelity, and Applicability (E/F/A). Experiments using various LLMs and research systems indicate that while retrieval tools can enhance evidence support, they do not reliably improve citation trustworthiness. The findings suggest that reliable legal research generation necessitates citation-aware governance, ensuring that retrieved legal authorities are not only found but also accurately described and appropriately applied. AI
IMPACT This benchmark could drive improvements in the reliability and trustworthiness of AI systems used in legal research.
RANK_REASON The item describes a new academic benchmark for evaluating AI systems in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
- arXivLabs
- Citation Trustworthiness
- Connected Papers
- Hugging Face
- LegalCiteTrust
- Litmaps
- python-coverage
- Retrieval tools
- scite Smart Citations
- Standard Chinese
- SUPPORT+
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →