A new benchmark has been developed to test how well Large Language Models (LLMs) ground their answers in actual Japanese legal articles. The benchmark uses the official e-Gov Law API v2, containing over 6,900 articles, to verify citations made by LLMs. When prompted in Japanese, two out of three tested local LLMs invented legal articles more frequently than when prompted in English, suggesting a language-dependent grounding issue. AI
IMPACT Highlights potential issues with LLM reliability in legal contexts and the importance of language-specific grounding.
RANK_REASON The item describes a new benchmark for evaluating LLM grounding in legal texts. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →