A new framework called CIVI has been developed to diagnose failures in search agents used for civic information, particularly those powered by large language models. The framework, which spans international government contexts and UN standards, found that none of the ten evaluated frontier search agents matched human baseline performance. CIVI also introduced ARISE, a method that attributes 72.1% of observed failures to retrieval issues rather than limitations in the models' core knowledge. AI
IMPACT This framework could improve the reliability of AI agents in public sector applications, reducing potential harm from incorrect information.
RANK_REASON The cluster contains an academic paper introducing a new framework and methodology for evaluating AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →