Researchers have developed ContractEval, a new diagnostic framework designed to evaluate the procedural instruction conformance of LLM agents. This system explicitly identifies when agents fail to meet specific obligations, such as skipping checks or violating invariants, which can lead to seemingly correct but unjustified outputs. ContractEval represents these procedural instructions as query-active obligations and matches them against response or trace evidence, distinguishing various types of conformance failures. In tests on audited procedural contracts, ContractEval successfully detected and localized all injected structural failures that were missed by traditional output-only and trace-aware judges. AI
IMPACT Enhances the auditability of LLM agent procedures, improving reliability for complex task execution.
RANK_REASON The cluster contains an academic paper detailing a new framework for evaluating LLM agents. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Connected Papers
- ContractEval
- CORE Recommender
- Hugging Face
- Litmaps
- LLM agents
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →