PulseAugur
中
实时 08:54:08
English(EN) RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

新的RL-ARC框架提高了LLM推理的置信度

研究人员推出RL-ARC,一个新颖的训练框架,旨在提高用于推理任务的大型语言模型(LLM)的校准并减少其过度自信。与以往忽略校准或牺牲推理性能的方法不同,RL-ARC联合优化推理置信度和答案置信度。该方法使用推理置信度作为信号,以惩罚错误答案的过度自信并规范正确答案,从而在不显著影响推理能力的情况下实现更可靠的置信度估计。 AI

影响 通过在不牺牲性能的情况下提高置信度估计,增强了推理模型的可靠性。

排序理由 该集群包含一篇详细介绍语言模型新训练框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的RL-ARC框架提高了LLM推理的置信度

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍语言模型新训练框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Gukhyeon Lee, SangKeun Lee ·

    RL-ARC:通过推理引导的不确定性校准大型推理模型

    arXiv:2610.11352v1 Announce Type: new Abstract: Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can l…