Researchers have developed a novel technique called Wiener Representation Filtering to reduce hallucinations in vision-language models (VLMs). This training-free method operates post-hoc by editing the representation space of the language backbone. By modeling hidden states as a combination of truthful and hallucinated components, the technique derives a Wiener-type estimator that attenuates hallucination-associated modes. Applied to models like LLaVA-1.5, MiniGPT-4, Gemma3, and mPLUG-Owl2, this method consistently lowers object hallucination on benchmarks such as CHAIR, POPE, and MME without affecting caption fluency or response quality. The approach also shows promise in video understanding and discrete diffusion language models. AI
IMPACT This technique could lead to more reliable and trustworthy vision-language models by reducing instances of fabricated information.
RANK_REASON The item is an academic paper detailing a new method for improving AI model performance. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Discrete Diffusion Language Models
- gemma3
- LLaVA-1.5
- MiniGPT-4
- mPLUG-Owl2
- POPE
- TempCompass
- vision-language model
- Wiener Representation Filtering
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →