PulseAugur
EN
LIVE 20:28:09

New VOPE benchmark reveals heavy hallucination in LVLMs during imagination tasks

A new evaluation benchmark called VOPE has been introduced to assess hallucinations in Large Vision-Language Models (LVLMs) during voluntary imagination tasks. Unlike previous research focusing on factual descriptions, VOPE specifically targets the models' ability to correctly judge the presence of imagined objects within their generated content. Experiments with current LVLMs and mitigation techniques reveal that most models exhibit significant hallucination issues in these imaginative scenarios, and existing methods are largely ineffective at addressing this problem. AI

IMPACT Highlights a critical gap in LVLM safety and reliability, particularly for creative and generative applications.

RANK_REASON Academic paper introducing a new benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New VOPE benchmark reveals heavy hallucination in LVLMs during imagination tasks

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Xingming Long, Jie Zhang, Shiguang Shan, Xilin Chen ·

    VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task

    arXiv:2511.13420v2 Announce Type: replace Abstract: Most research on hallucinations in Large Vision-Language Models (LVLMs) focuses on factual description tasks that prohibit any output absent from the image. However, little attention has been paid to hallucinations in voluntary …