Researchers have identified a "Representation-Action Gap" in omnimodal large language models, where models can encode mismatches between textual claims and their sensory input but fail to act on this information in their outputs. A new benchmark, IMAVB, was developed using movie clips to test this conflict detection capability. Across eight open-source models and Gemini 3.1 Pro, the study found that models either under-reject false claims or over-reject, impacting comprehension accuracy. The gap is more pronounced in audio than vision and is resistant to prompting, though a probe-guided logit adjustment technique showed promise in improving rejection behavior. AI
IMPACT Highlights a critical gap in LLM grounding, suggesting current models may misinterpret or ignore conflicting sensory data, impacting their reliability as agents.
RANK_REASON The cluster contains an academic paper detailing a new benchmark and findings about omnimodal LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →