A new benchmark has been developed to assess the reliability of large language models (LLMs) when applied to the Colombian legal system. The benchmark, which includes 1,042 items across ten legal areas and various question formats, revealed significant performance disparities among 15 evaluated models. While Gemini 3.1 Pro achieved high accuracy on closed-choice questions, factual correctness for free-text legal answers did not exceed 0.45 for any model. The study also highlighted a concerning dissociation between answer relevancy and correctness, indicating that LLMs often appear responsive but are factually inaccurate, necessitating expert supervision for legal tasks. AI
IMPACT Highlights the need for specialized benchmarks and expert oversight for LLMs in non-English legal systems, potentially impacting global adoption and trust.
RANK_REASON Academic paper introducing a new benchmark for LLM evaluation in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →