PulseAugur
实时 08:33:02
English(EN) Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

新框架GRAPHEVAL量化LLM推理的不确定性和连贯性

研究人员开发了GRAPHEVAL,一个新颖的基于图的框架,用于评估大型语言模型(LLM)的推理能力。该框架引入了图推理连贯性得分(GRCS)来量化LLM推理过程的语义和结构一致性,旨在检测诸如自信幻觉之类的问题。该研究还提出了图自洽性(GSC),这是一种解码策略,它优先考虑推理保真度而非原始准确性,特别是对于较小的模型,同时在能力更强的模型上保持或提高性能。 AI

影响 这项研究可能带来更可靠的LLM评估,推动模型具备更强大和更忠实的推理能力。

排序理由 该集群包含一篇学术论文,详细介绍了一个用于评估LLM推理的新框架和指标。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新框架GRAPHEVAL量化LLM推理的不确定性和连贯性

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Riccardo Revalor, Jalees Rehman, Debjit Pal ·

    我们能信任大语言模型的逻辑吗?通过基于图的框架量化不确定性、连贯性和鲁棒性

    arXiv:2607.08017v1 Announce Type: cross Abstract: Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of in…

  2. arXiv cs.CL TIER_1 English(EN) · Debjit Pal ·

    我们能信任大语言模型的逻辑吗?通过基于图的框架量化不确定性、连贯性和鲁棒性

    Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of intermediate steps. This raises three fundamental qu…