A new evaluation benchmark called VOPE has been introduced to assess hallucinations in Large Vision-Language Models (LVLMs) during voluntary imagination tasks. Unlike previous research focusing on factual descriptions, VOPE specifically targets the models' ability to correctly judge the presence of imagined objects within their generated content. Experiments with current LVLMs and mitigation techniques reveal that most models exhibit significant hallucination issues in these imaginative scenarios, and existing methods are largely ineffective at addressing this problem. AI
IMPACT Highlights a critical gap in LVLM safety and reliability, particularly for creative and generative applications.
RANK_REASON Academic paper introducing a new benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Large Vision Language Models
- ScienceCast
- Xingming Long
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →