研究人员正在开发检测和减轻大型语言模型(LLM)幻觉的新方法。一种方法涉及探测模型内部状态,以确定幻觉的确切发生和持续,表明外部观察者在检测方面与模型本身一样有效。另一种策略侧重于聚合多个廉价的开放权重模型,充当可靠的裁判,以较低的成本实现了接近前沿模型的性能。此外,针对特定模态(如视听LLM)的新技术正在涌现,通过引导内部问题状态来解决源混淆的接地幻觉。
AI
arXiv:2610.08026v1 Announce Type: new Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresp…
arXiv cs.AI
TIER_1English(EN)·Hyunjae Ra, Aecheon Jung, Jungin Park, Sungeun Hong·
arXiv:2610.02976v1 Announce Type: new Abstract: Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language model…
arXiv:2610.03482v1 Announce Type: new Abstract: Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt …
arXiv:2610.02066v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external …
arXiv:2601.19918v2 Announce Type: replace Abstract: Hallucinations in Large Language Models (LLMs), i.e., plausible but non-factual generations, pose a significant challenge to reliable deployment in high-stakes environments. However, many existing hallucination detectors require…
arXiv:2609.39182v1 Announce Type: cross Abstract: World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, …
arXiv cs.AI
TIER_1English(EN)·Elia Onofri, Roberto Di Pietro·
arXiv:2609.39229v1 Announce Type: cross Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models, costly and ill-suited to high-throughput monitoring. We investigat…
arXiv cs.CL
TIER_1English(EN)·Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman·
arXiv:2609.38355v1 Announce Type: new Abstract: We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. The shared task is organized into two main tasks, ea…
arXiv cs.AI
TIER_1English(EN)·Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon·
arXiv:2609.35804v1 Announce Type: cross Abstract: Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing co…
arXiv cs.AI
TIER_1English(EN)·Pranav Darshan, Pranav A, Sravan Karthick T, Minal Moharir, Ivan P. Yamshchikov·
arXiv:2609.35860v1 Announce Type: cross Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models a…
arXiv:2609.37568v1 Announce Type: new Abstract: Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs fa…
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confuse…
While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visu…
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: source-confused groundi…
arXiv:2609.38979v1 Announce Type: new Abstract: Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinat…
<!-- SC_OFF --><div class="md"><p>Hey everyone,</p> <p>If you run local models via Ollama in production or personal projects, you've probably run into the hallucination problem: how do you know when a model is hallucinating without burning extra VRAM or waiting 5 seconds for a he…