Researchers have developed C$^3$PO, a new benchmark designed to evaluate the cross-modal reasoning capabilities of omnimodal models. This benchmark, comprising 3,404 samples across video, audio, image, and text, specifically tests a model's ability to compose information from different modalities and resolve deliberate contradictions. The study found that while humans achieve high accuracy, even advanced models like Gemini-3.1 Pro struggle, with a significant portion of failures attributed to modality dominance where models prioritize one input type over others, particularly text. AI
IMPACT Highlights critical limitations in current omnimodal AI, suggesting future architectures must prioritize sustained cross-modal attention over modality dominance.
RANK_REASON The cluster contains a new academic paper detailing a novel benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →