PulseAugur
EN
LIVE 04:51:56

LLMs Gemma and Qwen struggle to detect hallucinations via logprobs

A Reddit user explored whether large language models like Gemma and Qwen can detect their own hallucinations by analyzing token probabilities (logprobs). The experiment suggested that when a model is uncertain, probabilities are distributed across multiple tokens, whereas a confident but incorrect recall might concentrate on a single wrong belief. The user developed a custom WebUI tool to test this, finding that both Gemma and Qwen performed poorly in utilizing their logprobs to identify uncertainty or hallucinations. AI

IMPACT Suggests current methods for LLMs to self-detect hallucinations are limited, impacting reliability.

RANK_REASON User-led exploration of LLM capabilities, not a formal research paper or product release.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLMs Gemma and Qwen struggle to detect hallucinations via logprobs

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Any-Chipmunk5480 ·

    Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vlvq2s/can_gemma_and_qwen_models_catch_hallucinations_by/"> <img alt="Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?" src="https://preview.redd.it/ohfin36yktih1.png?width=640…