A new evaluation pipeline called CERTID has been developed to assess the soundness of reasoning models in identifying causal effects from observational data. This pipeline addresses limitations in prior evaluations by using a formal identification algorithm and verifying formulas against structural causal models. When tested on three frontier models—Gemini Flash, Gemini Pro, and GPT5.5—CERTID revealed significant variations in their false-claim rates on non-identifiable queries, highlighting that accuracy is a poor indicator of soundness. AI
IMPACT Highlights critical limitations in current reasoning models' ability to reliably identify causal effects, suggesting a need for improved soundness in AI systems.
RANK_REASON The cluster describes a new research paper detailing a novel evaluation pipeline for AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →