PulseAugur
中
实时 13:00:49
English(EN) Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

新的训练方法教会LLM在不确定时弃权

研究人员开发了一种名为强化犹豫(RH)的新训练方法,通过教会语言模型在不确定时放弃回答,使其更加值得信赖。与奖励任何回答的传统方法不同,RH使用三元奖励,对错误回答的惩罚比弃权更严重。在逻辑谜题、医学问题和高等数学问题上的实验表明,RH能有效训练模型校准其诚实度,不同的惩罚级别可以产生针对不同风险容忍度的优化模型。该研究还引入了级联和自级联等推理策略,利用弃权作为协调信号,以更低的计算成本优于多数投票。 AI

影响 这项研究通过使AI系统能够准确地发出不确定信号,减少了在关键应用中幻觉的影响,有望带来更可靠、更值得信赖的AI系统。

排序理由 介绍语言模型新颖训练方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的训练方法教会LLM在不确定时弃权

本文如何被排名

Signal score
7 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
介绍语言模型新颖训练方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li ·

    诚实胜于准确:通过强化犹豫实现值得信赖的语言模型

    arXiv:2511.11500v3 Announce Type: replace Abstract: Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong a…