A new research paper published on arXiv has identified significant issues with the IntentBench benchmark, a key tool for evaluating audio-visual question-answering models. The researchers found that a substantial portion of the benchmark's questions are either broken or can be answered trivially without video input. They have released a cleaned version called Intentbench-Prime. Furthermore, the study suggests that current reasoning approaches for these models are costly and surprisingly ineffective, with a simple fine-tuned model (Vanilla SFT) performing comparably or better at a fraction of the cost. The research also indicates that models can learn significant social understanding priors solely from text, sometimes outperforming video-based methods. AI
IMPACT Highlights limitations in current multimodal LLMs for social understanding and suggests more efficient evaluation methods.
RANK_REASON Research paper published on arXiv detailing findings about an AI benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →