English(EN)Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models
新方法解决视频大模型中的幻觉问题
作者PulseAugur 编辑部·[4 个来源]·
研究人员开发了几种新方法来解决视频大模型(VLMMs)中的幻觉问题。一种方法 MultiToP,通过选择性地用全局补丁标记替换不可靠的视觉标记来在语言生成之前对其进行精炼。另一种方法 ViSSRes,使用轻量级网络增强视频表示,以提高时空和语义一致性。第三种技术侧重于精炼文本嵌入,以鼓励更好地整合视觉信息并减少对语言先验的过度依赖。这些方法在减少幻觉率和提高各种基准测试中的视频理解能力方面显示出显著的改进。
AI
arXiv:2606.11792v1 Announce Type: cross Abstract: Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully supported by the input video. In this paper, we propose …
Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully supported by the input video. In this paper, we propose MultiToP, a multimodal-context-aware visual token …
arXiv:2601.22574v2 Announce Type: replace-cross Abstract: Although Video Large Multimodal Models have achieved strong performance in video understanding, they still suffer from hallucination. Existing inference-time intervention methods usually modify videos under the contrastive…
arXiv:2511.05017v2 Announce Type: replace Abstract: Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual information during multimodal reasoning. A key cause is the model's over-reliance on text…