Researchers have developed CARVE, a new black-box counterfactual probe designed to audit the reliability of video agents. CARVE compares how an agent's answer changes when the semantic content of retrieved video frames is destroyed versus when the same pipeline is re-executed on those frames. In experiments with a VideoExplorer-style agent, CARVE demonstrated a significant and reproducible effect, altering the agent's answer more frequently when frame content was destroyed. The probe also showed potential as a routing signal, improving accuracy on a benchmark dataset. AI
IMPACT This new auditing method could improve the reliability and trustworthiness of AI agents that process visual information.
RANK_REASON The item is a research paper detailing a new auditing method for AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →