PulseAugur
实时 01:19:18
English(EN) LambdaPO: A Lambda Style Policy Optimization for Reasoning Language Models

AI 研究探索分层推理、反事实和高效训练方法 · 已追踪 10 个来源

几篇最新的研究论文探讨了 AI 推理和模型训练方面的先进技术。“Concept Flow Models” 引入了一种分层方法来提高基于概念的推理的可解释性,并减少信息泄露。“DeepSWIP” 为神经概率逻辑程序提出了一个反事实推理框架,增强了因果语义。“Vero” 提供了一个用于通用视觉推理的开放强化学习配方,旨在实现可复现性和可扩展性。此外,对“Reinforcement-aware Knowledge Distillation” 和 “Mechanism-Guided Selective Unlearning” 的研究解决了大型语言模型在推理任务中的训练和优化挑战,重点关注效率和防止能力回归。 AI

影响 这些论文通过提高模型的可解释性、因果推理、视觉推理能力以及大型语言模型的训练效率,推动了 AI 的发展。

排序理由 该集群包含多篇详细介绍新颖 AI 研究方法和发现的学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 421 个来源。 我们如何撰写摘要 →

AI 研究探索分层推理、反事实和高效训练方法 · 已追踪 10 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含多篇详细介绍新颖 AI 研究方法和发现的学术论文。
Source corroboration
421 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
319 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [421]

  1. arXiv cs.AI TIER_1 English(EN) · Gabriel Sarch, Linrong Cai, Qunzhong Wang, Haoyang Wu, Danqi Chen, Zhuang Liu ·

    Vero:一个用于通用视觉推理的开放式强化学习方法

    arXiv:2604.04917v3 Announce Type: replace-cross Abstract: What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) suggest that broad visual reasoning is within reach, …

  2. arXiv cs.AI TIER_1 English(EN) · Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, Stefano Soatto ·

    面向LLM推理的强化感知知识蒸馏

    arXiv:2602.22495v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller stud…

  3. arXiv cs.AI TIER_1 English(EN) · Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei ·

    超越推理能力:缓解大型推理模型中的通用能力遗忘

    arXiv:2510.21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language m…

  4. arXiv cs.AI TIER_1 English(EN) · Ya Wang, Adrian Paschke ·

    Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks

    arXiv:2606.19489v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space. Recent approaches leverage vision-language models to generate concept embeddings, reducing the nee…

  5. arXiv cs.AI TIER_1 English(EN) · Saimun Habib, Vaishak Belle, Fengxiang He ·

    DeepSWIP: 神经概率逻辑程序的商-WMC反事实

    arXiv:2606.20526v1 Announce Type: new Abstract: Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference is associational. Counterfactual reasoning additionally requires a causal semantics for interventions and evidence.…

  6. arXiv cs.AI TIER_1 English(EN) · Fengxiang He ·

    DeepSWIP:神经概率逻辑程序的商-WMC反事实

    Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference is associational. Counterfactual reasoning additionally requires a causal semantics for interventions and evidence. We introduce DeepSWIP, a single-world counterfa…

  7. arXiv cs.AI TIER_1 English(EN) · Gilad Yehudai, Clayton Sanford, Maya Bechler-Speicher, Orr Fischer, Ran Gilad-Bachrach, Amir Globerson ·

    Transformers 在图任务算法推理中的深度-宽度权衡

    arXiv:2503.01805v3 Announce Type: replace-cross Abstract: Transformers have revolutionized the field of machine learning. In particular, they can be used to solve complex algorithmic problems, including graph-based tasks. In such algorithmic tasks a key question is what is the mi…

  8. arXiv cs.AI TIER_1 English(EN) · Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou ·

    面向RLVR诱导推理的机制引导选择性遗忘

    arXiv:2606.19222v1 Announce Type: cross Abstract: We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially lower collateral damage than standard full-parameter updates. In matched SFT/RLVR checkpoi…

  9. arXiv cs.CL TIER_1 English(EN) · Zhuoran Li, Rui Xu, Jian Yang, Junnan Liu, Zhijun Chen, Qianren Mao, Hongcheng Guo, Jiaheng Liu, Likang Xiao, Ming Li, Xiaojie Wang ·

    通过可控模型合并增强多语言推理能力

    arXiv:2606.19002v1 Announce Type: new Abstract: Model merging is an effective technique for composing the capabilities of a multilingual model and a reasoning model. It has achieved promising generalization in multilingual reasoning tasks by aligning feature spaces of different m…

  10. arXiv cs.CL TIER_1 English(EN) · Yuliang Zhan, Xinyu Tang, Jian Li, Dandan Zheng, Weilong Chai, Jingdong Chen, Jun Zhou, Ge Wu, Wenyue Tang, Hao Sun ·

    GraphPO:基于图的策略优化用于推理模型

    arXiv:2606.18954v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses independently and optimizes the policy using from final an…

  11. arXiv cs.CL TIER_1 English(EN) · Jihyung Park, Minchao Huang, Leqi Liu, Elias Stengel-Eskin ·

    PragReST:用于实用语言理解的自强化反事实推理

    arXiv:2606.18624v1 Announce Type: new Abstract: Natural language understanding often depends on meanings that are implied rather than explicitly stated, requiring pragmatic reasoning. Despite strong performance on math and logical reasoning, large language models (LLMs) still str…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向RLVR诱导推理的机制引导选择性遗忘

    We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially lower collateral damage than standard full-parameter updates. In matched SFT/RLVR checkpoints on Qwen2.5-Math-1.5B and Qwen3-1.7B-Base, the …

  13. arXiv cs.AI TIER_1 English(EN) · Xu Zhou ·

    面向RLVR诱导推理的机制引导选择性遗忘

    We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially lower collateral damage than standard full-parameter updates. In matched SFT/RLVR checkpoints on Qwen2.5-Math-1.5B and Qwen3-1.7B-Base, the …

  14. arXiv cs.CL TIER_1 English(EN) · Xiaojie Wang ·

    通过可控模型合并增强多语言推理能力

    Model merging is an effective technique for composing the capabilities of a multilingual model and a reasoning model. It has achieved promising generalization in multilingual reasoning tasks by aligning feature spaces of different models. However, the merged single model often fa…

  15. arXiv cs.CL TIER_1 English(EN) · Hao Sun ·

    GraphPO:基于图的策略优化用于推理模型

    Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses independently and optimizes the policy using from final answers. This paradigm has two limitations. First,…

  16. arXiv cs.AI TIER_1 English(EN) · Bihao Zhan, Zongsheng Cao, Jie Zhou, Bo Zhang, Liang He ·

    FlowRAG:通过频率感知多粒度图流协同显式推理

    arXiv:2606.17856v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) is effective for knowledge-intensive and multi-hop query tasks; however, many existing methods primarily seed entity-based graphs and rely on implicit semantic relevance propagat…

  17. arXiv cs.AI TIER_1 English(EN) · Sajad Movahedi, Vera Milovanovi\'c, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto ·

    Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

    arXiv:2606.18206v1 Announce Type: new Abstract: Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solution these mo…

  18. arXiv cs.AI TIER_1 English(EN) · Baishali Chaudhury, Mengdie Flora Wang, Hyunji Hayley Park, Rahul Ghosh, Sungmin Hong, Jae Oh Woo ·

    通过结构化不确定性量化LLM逻辑推理的一致性

    arXiv:2606.17312v1 Announce Type: new Abstract: Large language models can arrive at the same answer through reasoning paths that are unstable, contradictory, or difficult to rank consistently -- a failure mode especially prevalent in multi-step deductive reasoning. Existing metho…

  19. arXiv cs.LG TIER_1 English(EN) · Chia-Hsuan Hsu, Jui-Ming Yao ·

    学习精炼隐藏状态以实现可靠的LLM推理

    arXiv:2606.17524v1 Announce Type: new Abstract: Large language models show strong reasoning ability, but their internal reasoning process can remain unstable in complex multi-step settings, where early hidden-state errors may propagate to incorrect predictions. We propose ReLAR, …

  20. arXiv cs.CL TIER_1 English(EN) · Jinyang Wu, Guocheng Zhai, Ruihan Jin, Jiahao Yuan, Yuhao Shen, Shuai Zhang, Zhengqi Wen, Jianhua Tao ·

    Atlas:为多领域复杂推理编排异构模型和工具

    arXiv:2601.03872v2 Announce Type: replace Abstract: The integration of large language models (LLMs) with external tools has significantly expanded the capabilities of AI agents. However, as the diversity of both LLMs and tools increases, selecting the optimal model-tool combinati…

  21. arXiv cs.CL TIER_1 English(EN) · Aryasomayajula Ram Bharadwaj ·

    通过闭环PID控制实现高效LLM推理的自适应激活引导

    arXiv:2506.18831v3 Announce Type: replace Abstract: Reasoning LLMs trained with long chain-of-thought often overthink: they spend tokens on redundant reflection and transitions that inflate cost without improving accuracy. Static activation steering (e.g.\ SEAL) suppresses such c…

  22. arXiv cs.CL TIER_1 English(EN) · Peixian Zhou, Yuxu Chen, Chaorui Zhang, Wei Han, Bo Bai, Xueyan Niu ·

    ChLogic:评估中文表达中逻辑推理的鲁棒性

    arXiv:2606.17905v1 Announce Type: new Abstract: Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogic, an English--Chinese aligned benchmark that tests …

  23. arXiv cs.CL TIER_1 English(EN) · Zihao Wei, Wenjie Shi, Liang Pang, Jingcheng Deng, Shicheng Xu, Shasha Guo, Zenghao Duan, Jiahao Liu, Jingang Wang, Huawei Shen, Xueqi Cheng ·

    动态部署编辑以减少RL训练推理模型的过度思考

    arXiv:2606.17890v1 Announce Type: new Abstract: Long-form chain-of-thought reasoning can improve LLM performance on complex tasks, but models often continue generating unnecessary reasoning after a correct answer has emerged. We refer to this behavior as overthinking. We study th…

  24. arXiv cs.AI TIER_1 English(EN) · Arshad Beg, Diarmuid O'Donoghue, Rosemary Monahan ·

    学习式形式推理:从合同合成到工件复用和形式语义

    arXiv:2602.02881v2 Announce Type: replace-cross Abstract: This paper articulates a long-term research vision for formal methods at the intersection with artificial intelligence, outlining multiple conceptual and technical dimensions and reporting on our ongoing work toward realis…

  25. arXiv cs.AI TIER_1 English(EN) · Jiahao Wang, Bingyu Liang, Chenhao Hu, Longhui Zhang, Xuebo Liu, Min zhang, Jing Li, Xuelong Li ·

    SuCo:充分性引导的连续自适应推理

    arXiv:2606.17687v1 Announce Type: cross Abstract: Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simple queries. Existing efforts to mitigate this ineff…

  26. arXiv cs.CL TIER_1 English(EN) · Elias Stengel-Eskin ·

    PragReST:用于实用语言理解的自强化反事实推理

    Natural language understanding often depends on meanings that are implied rather than explicitly stated, requiring pragmatic reasoning. Despite strong performance on math and logical reasoning, large language models (LLMs) still struggle with making pragmatic inferences, often ch…

  27. arXiv cs.AI TIER_1 English(EN) · Antonio Orvieto ·

    固定点推理器:稳定且自适应的深度循环 Transformer

    Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solution these models find. Like deep architectures, looped archi…

  28. arXiv cs.CL TIER_1 English(EN) · Xueyan Niu ·

    ChLogic:评估中文表达中逻辑推理的鲁棒性

    Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogic, an English--Chinese aligned benchmark that tests whether models preserve logical reasoning perfor…

  29. arXiv cs.CL TIER_1 English(EN) · Xueqi Cheng ·

    动态部署编辑以减少RL训练推理模型的过度思考

    Long-form chain-of-thought reasoning can improve LLM performance on complex tasks, but models often continue generating unnecessary reasoning after a correct answer has emerged. We refer to this behavior as overthinking. We study this phenomenon from the perspective of GRPO-style…

  30. arXiv cs.AI TIER_1 English(EN) · Liang He ·

    FlowRAG:通过频率感知多粒度图流协同显式推理

    Graph-based retrieval-augmented generation (GraphRAG) is effective for knowledge-intensive and multi-hop query tasks; however, many existing methods primarily seed entity-based graphs and rely on implicit semantic relevance propagation. This often (i) under-retrieves when user qu…

  31. arXiv cs.CL TIER_1 English(EN) · Xuelong Li ·

    SuCo:充分性引导的连续自适应推理

    Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simple queries. Existing efforts to mitigate this inefficiency typically rely on discrete reasoning modes…

  32. arXiv cs.LG TIER_1 English(EN) · Chuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan, Zijian Zhao, Zhengyu Chen, Yuchen Tian, Lijun Wu, Conghui He, Sirui Han, Yike Guo ·

    突破自然推理界限:形式逻辑验证的交错奖励

    arXiv:2601.22642v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid. To bridge this gap, we introduce a form…

  33. arXiv cs.LG TIER_1 English(EN) · David Huang, Lianlei Shan ·

    DLWM:用于高效多模态推理的多样化潜在世界模型

    arXiv:2606.15160v1 Announce Type: cross Abstract: Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years. Existing approaches typically rely on explicit chain-of-thought or continuous latent-space trajectories to enhance mult…

  34. arXiv cs.LG TIER_1 English(EN) · Lukas Fesser, Hanlin Zhang, Michelle M. Li, Eric Wang, Bryan Perozzi, Shekoofeh Azizi, Sham M. Kakade, Marinka Zitnik ·

    训练后如何塑造生物推理模型

    arXiv:2606.16517v1 Announce Type: new Abstract: Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are built through post-training, yet how each stage shapes …

  35. arXiv cs.LG TIER_1 English(EN) · Xian Sun, Wei Gao, Yingshuo Wang, Lingdong Kong, Yanhang Li, Zhichao Fan, Zexin Zhuang, Wenlong Dong, Zhiyuan Zheng, Hrishikesh Paranjape, Abhishek Mandal, Johnny R. Zhang ·

    超越准确性:在思维链推理中衡量偏见承认,以实现负责任的AI评估

    arXiv:2606.15127v1 Announce Type: new Abstract: Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support systems may require human oversight, and audit work…

  36. arXiv cs.CL TIER_1 English(EN) · Juming Xiong, Kevin Guo, Congning Ni, Wexin Liu, Chao Yan, Katherine Brown, Avinash Baidya, Xiang Gao, Bradley Malin, Zhijun Yin ·

    学习何时采样:置信度感知选择性采样以实现高效的思维链推理

    arXiv:2603.08999v3 Announce Type: replace Abstract: Large language models (LLMs) can achieve strong reasoning performance through chain-of-thought (CoT) reasoning, yet they often generate unnecessarily long reasoning paths that incur high inference cost. Self-consistency-based ap…

  37. arXiv cs.CL TIER_1 English(EN) · Jaehui Hwang, Byeongho Heo, Sangdoo Yun, Dongyoon Han ·

    哎呀,等等:论证模型中的语篇标记很重要

    arXiv:2601.17421v2 Announce Type: replace Abstract: Recent studies suggest that even data-efficient training with ($\simeq$1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through post-training. Such training corpora often contain …

  38. arXiv cs.CL TIER_1 English(EN) · Hoang Pham, Dong Le, Anh Tuan Luu ·

    GRACE:上下文忠实推理的步进基准

    arXiv:2606.16151v1 Announce Type: new Abstract: Many reasoning tasks require models to reason over input context, from document-grounded question answering to rule-based deduction. Chain-of-Thought (CoT) prompting produces traces that appear transparent, yet individual steps can …

  39. arXiv cs.CL TIER_1 English(EN) · Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin ·

    SciOrch:学习编排专家LLM以解决前沿多模态科学推理任务

    arXiv:2606.15872v1 Announce Type: new Abstract: Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest commercial systems fall short of expert-level performance. A closer look at model behavior reveals substantial comple…

  40. arXiv cs.CL TIER_1 English(EN) · Jiakai Li, Ke Qin, Rongzheng Wang, Yizhuo Ma, Qizhi Chen, Muquan Li, Shuang Liang ·

    当进一步推理无益时停止:推理模型中的注意力状态自适应生成

    arXiv:2606.15070v1 Announce Type: new Abstract: By incorporating test-time compute scaling, large reasoning models (LRMs) can solve complex problems through explicit chain-of-thought (CoT) reasoning processes. However, they often suffer from overthinking, resulting in redundant t…

  41. arXiv cs.CL TIER_1 English(EN) · Juming Xiong, Weixin Liu, Kevin Guo, Congning Ni, Junchao Zhu, Chongyu Qu, Chao Yan, Katherine Brown, Avinash Baidya, Xiang Gao, Bradley Malin, Zhijun Yin ·

    CoRA:置信度-理由对齐,实现可靠的思维链推理

    arXiv:2606.14961v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning can improve LLM performance, but high answer confidence may be misleading when the accompanying CoT rationale is plausible yet incomplete or poorly supported. We study confidence--rationale alignment…

  42. arXiv cs.AI TIER_1 English(EN) · Alex Bogdan ·

    自由能启发式:主动推理下的快速节俭认知与不确定精度

    arXiv:2606.15877v1 Announce Type: cross Abstract: Chain-of-thought (CoT) improves large language models' performance in math and symbolic reasoning. But on planning, contested ethics, and tasks where the model cannot check itself, more reasoning makes things worse. Both effects a…

  43. arXiv cs.AI TIER_1 English(EN) · Zhenyu Yu ·

    Vernier:探究因果推理中词汇差距背后的表征错位

    arXiv:2606.15733v1 Announce Type: cross Abstract: Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer ar…

  44. arXiv cs.AI TIER_1 English(EN) · Yu Li, Shu Hong, Tian Lan ·

    局部信用差异化:路径条件自蒸馏用于LLM推理

    arXiv:2606.15576v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy self-distillation addresses this by letting the same …

  45. arXiv cs.AI TIER_1 English(EN) · Dayeon Ki, Kevin Duh, Marine Carpuat ·

    AdaMame: 适应性多语言推理的训练配方

    arXiv:2606.15080v1 Announce Type: cross Abstract: While Large Reasoning Models (LRMs) show strong performance in English, they often fail to reason in the language of the query, a phenomenon known as language collapse. Existing RL-based fixes typically add a binary language fidel…

  46. arXiv cs.AI TIER_1 English(EN) · Keizo Kato, Chenhui Chu, Yugo Murawaki, Sado Kurohashi ·

    从最小标签扩展LLM推理:一种带有轻量级验证器的半监督框架

    arXiv:2606.16811v1 Announce Type: new Abstract: For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large numbers of correctly annotated answers to assess rea…

  47. arXiv cs.AI TIER_1 English(EN) · Ke Miao, Jiaxin Li, Hongliang Chen, Yuke Hu, Zhan Qin ·

    自适应与显式安全:触发大型推理模型中的潜在安全意识

    arXiv:2606.16808v1 Announce Type: new Abstract: While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data anno…

  48. arXiv cs.AI TIER_1 English(EN) · Yaoting Huang, Yifu Yuan, Linqi Han, Chengwen Li, Shuoheng Zhang, Xianze Yao, Hongyao Tang, Yan Zheng, Jianye Hao ·

    RoboPIN:通过固定思维链实现具身推理

    arXiv:2606.15753v1 Announce Type: new Abstract: Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models rely on text-…

  49. arXiv cs.AI TIER_1 English(EN) · Gowrav Mannem, Chowdhury Marzia Mahjabin, Jason Chen, Shivank Garg, Kevin Zhu ·

    序列模型在符号谜题上的递归推理

    arXiv:2606.15686v1 Announce Type: new Abstract: Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when problems become longer, harder, or slightly out of distribution. A major limitation of current r…

  50. Hugging Face Daily Papers TIER_1 English(EN) ·

    ChLogic:评估中文表达中逻辑推理的鲁棒性

    ChLogic benchmark reveals persistent performance gaps between English and Chinese logical reasoning in large language models, influenced by surface realization differences and translation artifacts.

  51. arXiv cs.AI TIER_1 English(EN) · Sado Kurohashi ·

    从最小标签扩展LLM推理:一种带有轻量级验证器的半监督框架

    For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large numbers of correctly annotated answers to assess reasoning quality. This paper presents a semi-super…

  52. arXiv cs.AI TIER_1 English(EN) · Zhan Qin ·

    自适应与显式安全:触发大型推理模型中的潜在安全意识

    While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe…

  53. arXiv cs.AI TIER_1 English(EN) · Alex Schutz, Victor-Alexandru Darvariu, Efimia Panagiotaki, Bruno Lacerda, Nick Hawes ·

    解决 GNARLy 问题:通过强化学习重塑图神经网络算法推理

    arXiv:2509.18930v3 Announce Type: replace-cross Abstract: Neural algorithmic reasoning (NAR) is a paradigm that trains neural networks to execute classic algorithms by supervised learning. Despite its successes, important limitations remain: inability to construct valid solutions…

  54. arXiv cs.AI TIER_1 English(EN) · Pratham Singla, Shivank Garg, Vihan Singh ·

    扑克竞技场:大型语言模型战略推理与记忆的多轴剖析

    arXiv:2606.13815v1 Announce Type: new Abstract: Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play benchmarks collapse heterogeneous reasoning dimensions into a single scalar, leaving the capabilit…

  55. arXiv cs.AI TIER_1 English(EN) · Zheyang Xiong, Shivam Garg, Max Yu, Vaishnavi Shrivastava, Haoyu Zhao, Anastasios Kyrillidis, Dimitris Papailiopoulos ·

    SuperThoughts:叠加态中的推理Token

    arXiv:2606.13862v1 Announce Type: cross Abstract: Long Chain-of-Thought (CoT) reasoning improves LLM problem-solving but is computationally expensive due to sequential token generation. While recent works explore reasoning in continuous latent spaces to bypass discrete token gene…

  56. arXiv cs.AI TIER_1 English(EN) · Avni Mittal, Rauno Arike ·

    C2-Faith:为链式思考推理中的因果和覆盖忠实度对 LLM 裁判进行基准测试

    arXiv:2603.05167v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as judges of chain-of-thought (CoT) reasoning, yet it remains unclear whether they can reliably assess process faithfulness rather than merely answer plausibility. We intr…

  57. arXiv cs.AI TIER_1 English(EN) · Dake Bu, Wei Huang, Andi Han, Atsushi Nitanda, Bo Xue, Qingfu Zhang, Hau-San Wong, Taiji Suzuki ·

    训练后分布偏差:推理轨迹的马尔可夫分析

    arXiv:2511.07368v3 Announce Type: replace-cross Abstract: Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiable rewards (RLVR) and test-time scaling (TTS). While recent work highlights the rol…

  58. arXiv cs.CL TIER_1 English(EN) · Anh Tuan Luu ·

    GRACE:上下文忠实推理的步进基准

    Many reasoning tasks require models to reason over input context, from document-grounded question answering to rule-based deduction. Chain-of-Thought (CoT) prompting produces traces that appear transparent, yet individual steps can silently deviate from the source evidence, even …

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    SciOrch:学习编排专家LLM以解决前沿多模态科学推理任务

    SciOrch is a framework that uses a lightweight orchestrator model to coordinate multiple frontier LLMs for scientific reasoning, achieving superior performance through MCTS-based training and GRPO-style optimization while reducing API costs.

  60. arXiv cs.CL TIER_1 English(EN) · Darpan Aswal, Thomas Palmeira Ferraz, Yongxin Zhou, Maxime Peyrard ·

    可观察模式并非解释:潜在推理模型的因果几何分析

    arXiv:2606.12689v1 Announce Type: new Abstract: Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decodable arithmetic computation, as evidence for interna…

  61. arXiv cs.AI TIER_1 English(EN) · Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Rapha\"el Milli\`ere, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Do… ·

    MoReBench:评估语言模型中的程序性和多元化道德推理,超越结果

    arXiv:2510.16380v2 Announce Type: replace-cross Abstract: As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but al…

  62. arXiv cs.AI TIER_1 English(EN) · Zilin Xiao, Qi Ma, Chun-cheng Jason Chen, Xintao Chen, Avinash Atreya, Hanjie Chen, Vicente Ordonez ·

    通过检索增强强化微调学习类比推理

    arXiv:2606.13680v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic similarity is poorly suited for complex reasoning ta…

  63. arXiv cs.AI TIER_1 English(EN) · Akshay Krishnamurthy, Audrey Huang, Nived Rajaraman ·

    选择与改进:理解推理的训练后机制

    arXiv:2606.13125v1 Announce Type: cross Abstract: Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capa…

  64. arXiv cs.AI TIER_1 English(EN) · Daniel Scalena, Sara Candussio, Luca Bortolussi, Elisabetta Fersini, Malvina Nissim, Gabriele Sarti ·

    超越承诺边界:探究大型推理模型中的表观因果链思维

    arXiv:2606.13603v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning is the dominant paradigm for inference-time scaling in language models, yet the causal influence of individual steps on the final answer poorly understood. We estimate each step's causal importance…

  65. arXiv cs.AI TIER_1 English(EN) · Sarah Elshabrawy, Rahul K. Dass, Ashok K. Goel ·

    构建程序推理的评估数据集:平衡自然性、接地性和多跳覆盖率

    arXiv:2606.12767v1 Announce Type: new Abstract: Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use. We study how TMK-based question…

  66. arXiv cs.AI TIER_1 English(EN) · Pierre Beckmann, Marco Valentino, Andre Freitas ·

    SciR:LLM 科学推理的可控基准

    arXiv:2606.13020v1 Announce Type: new Abstract: Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on …

  67. arXiv cs.AI TIER_1 English(EN) · Xin Wang, Boyan Gao, Yibo Yang, David A. Clifton ·

    Mental-R1: 对齐大语言模型推理以进行心理健康评估

    arXiv:2606.13176v1 Announce Type: new Abstract: Mental health problems such as anxiety, depression, and suicide remain urgent global challenges, where timely and accurate assessment is critical for effective intervention. Recently, large language models have been explored for men…

  68. arXiv cs.AI TIER_1 English(EN) · Fabrizio Marozzo, Pietro Li\`o ·

    LLM作为调查员:基于证据的推理用于鲁棒的交互式问题诊断

    arXiv:2606.13220v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving. However, when users provide incomplete descriptions or plausible but unverified explanations, LLMs may prematurely align wit…

  69. arXiv cs.AI TIER_1 English(EN) · Zach Studdiford, Gary Lupyan ·

    推理即模式匹配:人类与大型语言模型日常推理的共享机制

    arXiv:2606.13607v1 Announce Type: new Abstract: When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that p…

  70. arXiv cs.CL TIER_1 English(EN) · Yaniv Nikankin, Martin Tutek, Tomer Ashuach, Jonathan Rosenfeld, Yonatan Belinkov ·

    推理模型知道什么重要,并将其编码在它们的激活中

    arXiv:2604.18307v2 Announce Type: replace Abstract: Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the final answer, others are removable. Determining whi…

  71. arXiv cs.CL TIER_1 English(EN) · Nathaniel Bottman, Yinhong Liu, Kyle Richardson ·

    算子一致性:大型语言模型中组合推理失败的无标签信号

    arXiv:2606.13649v1 Announce Type: new Abstract: Detecting LLM reasoning failures at inference time without ground-truth labels has motivated a wide range of confidence baselines, including self-consistency, semantic entropy, and P(True), built on within-question sampling and self…

  72. arXiv cs.CL TIER_1 English(EN) · Nathaniel Bottman, Kyle Richardson ·

    用于大型语言模型组合推理的算子

    arXiv:2606.13634v1 Announce Type: new Abstract: Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM reasoning, yet it currently lacks a rigorous mathematic…

  73. arXiv cs.CL TIER_1 English(EN) · Shu Tong Luo, Wenqin Liu, Rui Liu, Mingming Gong, Jiaxian Guo ·

    分块到达上下文时的多轮推理:可扩展分片与增强记忆的强化学习

    arXiv:2606.12941v1 Announce Type: new Abstract: When a user reveals task-critical information across several conversation turns, LLM accuracy drops by up to 65% despite full context availability. We show that this Lost in Conversation degradation can be substantially mitigated by…

  74. arXiv cs.CL TIER_1 English(EN) · Dimitris Papailiopoulos ·

    SuperThoughts:叠加态中的推理Token

    Long Chain-of-Thought (CoT) reasoning improves LLM problem-solving but is computationally expensive due to sequential token generation. While recent works explore reasoning in continuous latent spaces to bypass discrete token generation, they often struggle with training stabilit…

  75. arXiv cs.CL TIER_1 English(EN) · Vihan Singh ·

    扑克竞技场:大型语言模型战略推理与记忆的多轴剖析

    Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play benchmarks collapse heterogeneous reasoning dimensions into a single scalar, leaving the capability structure of frontier LLMs unexamined. We intr…

  76. arXiv cs.AI TIER_1 English(EN) · Vicente Ordonez ·

    通过检索增强强化微调学习类比推理

    Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic similarity is poorly suited for complex reasoning tasks: a semantically similar problem may demand an …

  77. arXiv cs.CL TIER_1 English(EN) · Kyle Richardson ·

    算子一致性:大型语言模型组合推理失败的无标签信号

    Detecting LLM reasoning failures at inference time without ground-truth labels has motivated a wide range of confidence baselines, including self-consistency, semantic entropy, and P(True), built on within-question sampling and self-evaluation. Operad theory, the formalism for sy…

  78. arXiv cs.CL TIER_1 English(EN) · Kyle Richardson ·

    用于大型语言模型中组合推理的算子

    Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM reasoning, yet it currently lacks a rigorous mathematical foundation. In this paper, we propose operads…

  79. arXiv cs.AI TIER_1 English(EN) · Gary Lupyan ·

    推理即模式匹配:人类与大型语言模型日常推理的共享机制

    When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types…

  80. Hugging Face Daily Papers TIER_1 English(EN) ·

    推理即模式匹配:人类与大型语言模型日常推理的共享机制

    When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types…

  81. arXiv cs.AI TIER_1 English(EN) · Gabriele Sarti ·

    超越承诺边界:探究大型推理模型中的表观因果思维链

    Chain-of-thought (CoT) reasoning is the dominant paradigm for inference-time scaling in language models, yet the causal influence of individual steps on the final answer poorly understood. We estimate each step's causal importance via early exit and use this measure to study how …

  82. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Pietro Liò ·

    LLM作为调查员:基于证据的推理用于鲁棒的交互式问题诊断

    Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving. However, when users provide incomplete descriptions or plausible but unverified explanations, LLMs may prematurely align with these assumptions and propose solutions before…

  83. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM作为调查员:基于证据的推理,用于鲁棒的交互式问题诊断

    Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving. However, when users provide incomplete descriptions or plausible but unverified explanations, LLMs may prematurely align with these assumptions and propose solutions before…

  84. arXiv cs.LG TIER_1 English(EN) · Nived Rajaraman ·

    选择与改进:理解推理的训练后机制

    Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capabilities are acquired or enhanced via reinforcemen…

  85. arXiv cs.CL TIER_1 English(EN) · Jiaxian Guo ·

    分块到达上下文时的多轮推理:可扩展分片与增强记忆的强化学习

    When a user reveals task-critical information across several conversation turns, LLM accuracy drops by up to 65% despite full context availability. We show that this Lost in Conversation degradation can be substantially mitigated by training models to maintain a compact rolling m…

  86. arXiv cs.AI TIER_1 English(EN) · Valentin No\"el ·

    推理的几何:有效数学推理的光谱特征

    arXiv:2601.00791v2 Announce Type: replace-cross Abstract: Verifying whether a language model is genuinely reasoning or pattern-matching remains an open problem: learned verifiers are expensive, and output-based heuristics are brittle. We show that valid mathematical reasoning ind…

  87. arXiv cs.AI TIER_1 English(EN) · Jana Zeller, Thadd\"aus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, Wieland Brendel ·

    MentisOculi:揭示心像推理的局限性

    arXiv:2602.02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation. This shift has sparked interest …

  88. arXiv cs.AI TIER_1 English(EN) · Chao Lei, Guang Hu, Meng Yang, Yanbei Jiang, Nir Lipovetzky ·

    注意视角:让我们进行递归推理以实现心智理论

    arXiv:2606.11724v1 Announce Type: new Abstract: Theory of Mind (ToM) reasoning requires inferring agents' beliefs from partial and asymmetric observations, which remains an open challenge for LLMs. Existing prompting-based approaches improve ToM reasoning through observable-event…

  89. arXiv cs.AI TIER_1 English(EN) · Rikard Rosenbacke, Carl Rosenbacke, Victor Rosenbacke, Martin McKee ·

    从消费到反思:设计人机关系以实现稳定推理

    arXiv:2606.11195v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed how humans access information, but not how we reason with it. Their fluency accelerates consumption while bypassing the slow, reflective processes that underpin sound judgment. This pa…

  90. arXiv cs.AI TIER_1 English(EN) · Prakul Sunil Hiremath, Harshit R. Hiremath ·

    推理中的校准漂移:思维链预算如何在大语言模型中诱导过度自信

    arXiv:2606.11211v1 Announce Type: cross Abstract: The ability of large language models (LLMs) to express calibrated uncertainty is important for safe deployment. Chain-of-thought (CoT) reasoning is widely used to improve accuracy and reliability, but its effect on calibration is …

  91. arXiv cs.AI TIER_1 English(EN) · Subbarao Kambhampati, Karthik Valmeekam, Siddhant Bhambri, Vardhan Palod, Lucas Saldyt, Kaya Stechly, Soumya Rani Samineni, Durgesh Kalwar, Upasana Biswas ·

    立场:停止将中间标记拟人化为推理/思考痕迹!

    arXiv:2504.09762v4 Announce Type: replace Abstract: Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks. These intermediate tokens have been called \s…

  92. arXiv cs.AI TIER_1 English(EN) · Jiahao Yu, Zelei Cheng, Xian Wu, Xinyu Xing ·

    GPO:从关键步骤中学习以改进LLM推理

    arXiv:2509.16456v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in various domains, showing impressive potential on different tasks. Recently, reasoning LLMs have been proposed to improve the \textit{reasoning} or \textit{thinking} capabilit…

  93. arXiv cs.CL TIER_1 English(EN) · Hao Xiang, Qiaoyu Tang, Le Yu, Yaojie Lu, Xianpei Han, Ben He, Le Sun, Bowen Yu, Peng Wang, Hongyu Lin, Dayiheng Liu ·

    可验证环境是乐高积木:递归组合实现推理泛化

    arXiv:2606.12373v1 Announce Type: new Abstract: Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language Models (LLMs). While prior research demonstrates that scaling environment quantit…

  94. arXiv cs.CL TIER_1 English(EN) · Yijie Deng, He Zhu, Wen Wang, Junyou Su, Minxin Chen, Wenjia Zhang ·

    人工智能能像城市规划师一样进行推理吗?将大型语言模型与专业判断进行基准测试

    arXiv:2606.11678v1 Announce Type: new Abstract: Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key question for urban planning: which forms of professional planning knowledge can AI replicate, and which still require human judgment? Al…

  95. arXiv cs.CL TIER_1 English(EN) · Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar, Erik Cambria, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah ·

    大语言模型推理的周期表:推理范式、方法和失败模式的结构化调查

    arXiv:2606.11470v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved strong performance across natural language processing tasks, yet reliable reasoning remains an open challenge. Although modern LLMs show progress in structured inference, multi-step problem…

  96. arXiv cs.LG TIER_1 English(EN) · Hongyi Liu, Frederic Sala, Thomas Reps, Adithya Murali ·

    使用推理代理进行大规模反例引导学习

    arXiv:2606.11521v1 Announce Type: new Abstract: LLMs and LLM agents should improve when given feedback, but identifying when they are able to do so is difficult: feedback is heterogeneous, domain-specific, and difficult to control. We approach this challenge by asking LLMs to per…

  97. arXiv cs.CL TIER_1 English(EN) · Dayiheng Liu ·

    可验证环境是乐高积木:递归组合用于推理泛化

    Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language Models (LLMs). While prior research demonstrates that scaling environment quantity improves RL performance, existing manual or in…

  98. arXiv cs.CL TIER_1 English(EN) · Wenjia Zhang ·

    人工智能能否像城市规划师一样进行推理?将大型语言模型与专业判断进行基准测试

    Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key question for urban planning: which forms of professional planning knowledge can AI replicate, and which still require human judgment? Although AI tools are increasingly used in plannin…

  99. Hugging Face Daily Papers TIER_1 English(EN) ·

    人工智能能否像城市规划师一样进行推理?将大型语言模型与专业判断进行基准测试

    Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key question for urban planning: which forms of professional planning knowledge can AI replicate, and which still require human judgment? Although AI tools are increasingly used in plannin…

  100. arXiv cs.LG TIER_1 English(EN) · Evgenii Kortukov, Piotr Komorowski, Florian Klein, Paula Engl, Gabriele Sarti, Seong Joon Oh, Sebastian Lapuschkin, Wojciech Samek ·

    预测推理模型中的未来行为可实现更好的引导

    arXiv:2606.11172v1 Announce Type: new Abstract: Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output quality. We argue that prior steering work implicitl…

  101. arXiv cs.CL TIER_1 English(EN) · Adi Gabay, Gabriel Stanovsky, Liat Peterfreund ·

    超越记忆:使用认知谜题区分大型语言模型中的基于模式的推理和认知推理

    arXiv:2603.21350v2 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior work evaluating LLMs on epistemic puzzles often frames failures as memorization r…

  102. arXiv cs.CL TIER_1 English(EN) · Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata ·

    轻量级潜在推理用于叙事任务

    arXiv:2512.02240v2 Announce Type: replace Abstract: Large language models (LLMs) tackle complex tasks by generating long chains of thought or "reasoning traces" that act as latent variables in the generation of an output given a query. A model's ability to generate such traces ca…

  103. arXiv cs.CL TIER_1 English(EN) · Zhichen Dong, Yang Li, Yuhan Sun, Weixun Wang, Yijia Luo, Zinian Peng, Taiheng Ye, Chao Yang, Wenbo Su, Yu Cheng, Bo Zheng, Junchi Yan ·

    推理如何流动?追踪注意力诱导的信息流以实现 LLM 中的目标 RL

    arXiv:2606.10646v1 Announce Type: cross Abstract: Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing to distinguish decisive reasoning steps from routin…

  104. arXiv cs.CL TIER_1 English(EN) · Prajakta Kini, Avinash Reddy, Souradip Chakraborty, Satya Sai Srinath Namburi GNVV, Furong Huang, Amrit Singh Bedi, Alvaro Velasquez ·

    推理是否能保持对齐?论大型推理模型的可靠性

    arXiv:2606.11046v1 Announce Type: new Abstract: Instruction-tuned LLMs are increasingly converted into reasoning models through post-training to improve multi-step task performance. This conversion is usually optimized for reasoning accuracy, without explicitly preserving the ali…

  105. arXiv cs.CL TIER_1 English(EN) · Sanghee Park, Geewook Kim, Kee-Eung Kim ·

    KCSAT-ML:用全国性队列人类难度探测推理模型

    arXiv:2606.10403v1 Announce Type: new Abstract: Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean College Scholastic Ability Test (KCSAT; Suneung) mat…

  106. arXiv cs.AI TIER_1 English(EN) · Daeyong Kwon, Soyoung Yoon, Seung-won Hwang ·

    SAFE:一个 LLM 作为验证器框架,用于基于证据的多跳推理

    arXiv:2604.01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid intermediate reasoning. We propose SAFE, an LLM-as-verifier framework for evide…

  107. arXiv cs.AI TIER_1 English(EN) · Yubo Li, Lu Zhang, Tianchong Jiang, Ramayya Krishnan, Rema Padman ·

    模型称“走”:表面启发式方法如何压倒大型语言模型推理中的隐式约束

    arXiv:2603.29025v3 Announce Type: replace-cross Abstract: Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint. We introduce the Heuristic Override Benchmark (HOB): 500 instances spanning 4 heuristic families and 5 constraint fam…

  108. arXiv cs.AI TIER_1 English(EN) · Daniel Herbst, Lea Karbevska, Divyanshu Kumar, Akanksha Ahuja, Fatemeh Gholamzadeh Nasrabadi, Fabrizio Frasca ·

    序列化中的迷失:LLM图推理器的不变性和泛化性

    arXiv:2511.10234v3 Announce Type: replace-cross Abstract: While promising, graph reasoners based on Large Language Models (LLMs) lack built-in invariance to symmetries in graph representations. Operating on sequential graph serializations, LLMs can produce different outputs under…

  109. arXiv cs.AI TIER_1 English(EN) · Wooil Jung ·

    Dropout-GRPO:用于连续潜在推理的变分随机性

    arXiv:2606.10184v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) relies on the diversity of $K$ rollouts within each group; otherwise, the group-mean advantage $A^{(k)} = r^{(k)} - \mu_r$ collapses to zero. This presents a structural challenge for laten…

  110. arXiv cs.AI TIER_1 English(EN) · Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi ·

    当思维链更胜一筹时:多轮推理模型的失效模式

    arXiv:2606.10740v1 Announce Type: new Abstract: Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refusal rate may appear indistinguishable from a robustl…

  111. arXiv cs.AI TIER_1 English(EN) · Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li, Jianlong Chen, Xiangfeng Wang ·

    RealMath-Eval:为何 SOTA 评判模型难以理解真实人类推理

    arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined. To …

  112. Hugging Face Daily Papers TIER_1 English(EN) ·

    可验证环境是乐高积木:递归组合用于推理泛化

    Recursive automated composition framework enables scalable reinforcement learning for language models by automatically combining verifiable environments through compositional operators.

  113. arXiv cs.LG TIER_1 English(EN) · Wojciech Samek ·

    预测推理模型中的未来行为可实现更好的引导

    Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output quality. We argue that prior steering work implicitly relies on internal features that detect behavi…

  114. arXiv cs.CL TIER_1 English(EN) · Alvaro Velasquez ·

    推理是否能保持对齐?论大型推理模型的可靠性

    Instruction-tuned LLMs are increasingly converted into reasoning models through post-training to improve multi-step task performance. This conversion is usually optimized for reasoning accuracy, without explicitly preserving the alignment behavior of the instruction-tuned model, …

  115. Hugging Face Daily Papers TIER_1 English(EN) ·

    当思维链更胜一筹时:多轮推理模型的失效模式

    Multi-turn reasoning models exhibit hidden alignment failures that are masked by traditional evaluation methods, revealing vulnerabilities through a trace-level diagnostic framework that identifies distinct failure modes including context-injection failures.

  116. arXiv cs.AI TIER_1 English(EN) · Samuele Poppi ·

    当思维链更胜一筹时:多轮推理模型的失效模式

    Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refusal rate may appear indistinguishable from a robustly aligned baseline. To expose these hidden tempo…

  117. arXiv cs.CL TIER_1 English(EN) · Junchi Yan ·

    推理如何流动?追踪注意力诱导的信息流以实现 LLM 中的定向 RL

    Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing to distinguish decisive reasoning steps from routine formatting or fluent filler. Recent attempts lev…

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    推理如何流动?追踪注意力诱导的信息流以实现LLM中的定向RL

    FlowTracer is an RL framework that uses attention-induced graphs to trace reasoning flows and assign token-level credit based on global information propagation structures.

  119. arXiv cs.CL TIER_1 English(EN) · Kee-Eung Kim ·

    KCSAT-ML:用全国性队列人类难度探测推理模型

    Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean College Scholastic Ability Test (KCSAT; Suneung) mathematics: 664 problems with a 339-item core set …

  120. arXiv cs.AI TIER_1 English(EN) · Sanjay Kariyappa, G. Edward Suh ·

    指令层级失效:诊断与修复推理语言模型的失败

    arXiv:2606.07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the model should obey the highest-privilege applicable instruction. Existing benchmarks…

  121. arXiv cs.AI TIER_1 English(EN) · Mujtaba Farhan, Maheep Chaudhary ·

    为何将残差流限制在层而非token?用于连续潜在推理的持久内存

    arXiv:2606.07720v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks. The CoCoNuT (Chain of Continuous Thought) paradigm~\cite{hao2024coconut} extends this by enabling models to …

  122. arXiv cs.LG TIER_1 English(EN) · Yang Li, Zhichen Dong, Yuhan Sun, Weixun Wang, Shaopan Xiong, Yijia Luo, Jiashun Liu, Han Lu, Jiamang Wang, Wenbo Su, Bo Zheng, Junchi Yan ·

    注意力机制照亮LLM推理:预规划与锚定节奏实现细粒度策略优化

    arXiv:2510.13554v2 Announce Type: replace-cross Abstract: The reasoning pattern of Large language models (LLMs) remains opaque, and reinforcement learning (RL) typically applies uniform credit across an entire generation, blurring the distinction between pivotal and routine steps…

  123. arXiv cs.LG TIER_1 English(EN) · Zhanke Zhou, Xiangyu Lu, Chentao Cao, Brando Miranda, Tongliang Liu, Bo Han, Sanmi Koyejo ·

    易、难与可学:LLM推理的置信度与难度自适应策略优化

    arXiv:2606.07950v1 Announce Type: new Abstract: RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling and weighting, leading to inefficient compute alloc…

  124. arXiv cs.AI TIER_1 English(EN) · Xiaoou Liu, Tiejin Chen, Dengjia Zhang, Yaqing Wang, Lu Cheng, Hua Wei ·

    通过分步置信度归因诊断黑盒大模型的多步推理失败

    arXiv:2605.19228v2 Announce Type: replace-cross Abstract: Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning trace might fail remains difficult. Confide…

  125. arXiv cs.AI TIER_1 English(EN) · Javier Mar\'in ·

    Transformer 如何拒绝错误答案:事实约束处理的旋转动力学

    arXiv:2603.13259v2 Announce Type: replace-cross Abstract: When a decoder-only transformer is forced to process matched correct and incorrect single-token continuations of a factual query, the two pathways through hidden-state space diverge in a specific way: displacement vectors …

  126. arXiv cs.AI TIER_1 English(EN) · Shivam Adarsh, Maria Maistro, Christina Lioma ·

    语境如何塑造真相:LLM中语句级真值表征的几何变换

    arXiv:2601.06599v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studied in prior work, however how they change w…

  127. arXiv cs.AI TIER_1 English(EN) · Onat Ozer, Yuchen Wang, Grace Wu, Daniel Dosti, Honghao Zhang, Vivi De La Rue ·

    MAR:多智能体反思提升大型语言模型推理能力

    arXiv:2512.20845v2 Announce Type: replace Abstract: LLMs have shown the capacity to improve their performance on reasoning tasks through reflecting on their mistakes, and acting with these reflections in mind. However, continual reflections of the same LLM onto itself exhibit deg…

  128. arXiv cs.AI TIER_1 English(EN) · Junkai Zhang, Jingru Gan, Xiaoxuan Wang, Zian Jia, Changquan Gu, Jianpeng Chen, Yanqiao Zhu, Mingyu Derek Ma, Dawei Zhou, Ling Li, Wei Wang ·

    MatSciBench:对大语言模型在材料科学领域推理能力进行基准测试

    arXiv:2510.12171v2 Announce Type: replace Abstract: Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To fill this gap, we introduce MatSciBench, a comprehensive college-level benchmark …

  129. arXiv cs.AI TIER_1 English(EN) · Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth ·

    LLM驱动的解释实现健全且完整的神经符号推理

    arXiv:2507.09751v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but exhibit problems with logical consistency in their output. How can we harness LLMs' broad-coverage para…

  130. arXiv cs.AI TIER_1 English(EN) · Subramanyam Sahoo ·

    用于诊断推理模型中未知未知信息的结构化无知证书的校准

    arXiv:2606.08571v1 Announce Type: cross Abstract: Large language models frequently fail in a characteristic way: rather than acknowledging ignorance, they produce fluent but incorrect answers to questions that lie beyond their knowledge boundaries. We introduce \textbf{Structured…

  131. arXiv cs.AI TIER_1 English(EN) · Hengxin Fan ·

    能力而非格式:重新思考结构化推理的失败

    arXiv:2606.09410v1 Announce Type: new Abstract: Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model's spare capacity. Using information-matched prose controls and a four-level schema complexity…

  132. arXiv cs.AI TIER_1 English(EN) · Xinyue Liang, Yizhe Yang, Yu Bai, Bin Xu, Jiawei Li, Yang Gao ·

    不同的思维模式能提升大型语言模型的推理能力

    arXiv:2606.08974v1 Announce Type: new Abstract: Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reasoning chains. In this work, we focus on two critical yet underexplored aspects of …

  133. arXiv cs.AI TIER_1 English(EN) · Syed Rifat Raiyan, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan ·

    人工智能在数学推理中的应用:语言模型、神经符号系统与验证发现的综合调查

    arXiv:2606.08728v1 Announce Type: new Abstract: Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified acc…

  134. arXiv cs.AI TIER_1 English(EN) · Beiwen Zhang, Yongheng Liang, Guowei Zou, Haitao Wang, Hejun Wu ·

    将大型语言模型的推理提炼为可解释的策略树,以实现人机协作

    arXiv:2606.08596v1 Announce Type: new Abstract: Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of work. Most prior work relies on multi-agent reinforcement learning (MARL) to lear…

  135. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Foutse Khomh ·

    用于LLM鲁棒上下文推理的博弈论多智能体控制

    Large Language Models (LLMs) in multi-turn interactions maintain evolving context rather than generating isolated responses, making them vulnerable to prompt-injection and context-poisoning attacks in which locally plausible adversarial fragments gradually distort reasoning traje…

  136. arXiv cs.CL TIER_1 English(EN) · André Freitas ·

    无黄金标准推理:自动形式化的代理裁判理论

    Complex reasoning tasks increasingly require systems to produce outputs whose correctness cannot be judged by exact match against a single reference. Autoformalization (AF) is a representative example; it asks a model to translate informal mathematical or logical reasoning into a…

  137. arXiv cs.AI TIER_1 English(EN) · Hengxin Fan ·

    能力而非格式:重新思考结构化推理的失败

    Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model's spare capacity. Using information-matched prose controls and a four-level schema complexity gradient, we separate format-specific effects f…

  138. arXiv cs.CL TIER_1 English(EN) · Christina Niklaus ·

    TruthSplit:通过多视角推理实现论证中的条件有效性操作化

    We present TruthSplit, an interactive system for multi-perspective argument analysis. Existing argumentation tools typically analyze properties of the argument itself, such as structure, quality, stance, or persuasiveness, while leaving perspective-specific background knowledge i…

  139. arXiv cs.CL TIER_1 English(EN) · Huajun Chen ·

    使用复杂视觉查询进行符号和抽象推理

    Understanding and reasoning over abstract visual content remains a challenge for current multi-modal large language models (MLLMs). In this paper, we explore a novel abstract data type termed complex visual query (CVQ), designed to probe symbolic and abstractive reasoning, which …

  140. arXiv cs.CL TIER_1 English(EN) · Liang Wang ·

    CRANE:面向推理多模态大模型的知识编辑

    The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought (CoT) reasoning before producing answers, has introduced a new challenge for knowledge editing: methods that appear successful under traditional metrics (teacher-forcing …

  141. arXiv cs.AI TIER_1 English(EN) · Tanvi Thoria, Kiana Jafari, Marc R. Schlichting, Mykel J. Kochenderfer ·

    语言模型如何失效:已承诺和持续推理失败的 Token 级签名

    arXiv:2606.06635v1 Announce Type: cross Abstract: Failures in language model reasoning emerge through distinct processes that leave identifiable signatures in the reasoning trace. We characterize these failures using token-level uncertainty signals, finding they arise through two…

  142. arXiv cs.AI TIER_1 English(EN) · Vladislav Smirnov (MBZUAI), Chieu Nguyen (MBZUAI), Sergey Senichev (Independent Researcher), Minh Ngoc Ta (MBZUAI), Ekaterina Fadeeva (ETH Z\"urich), Artem Vazhentsev (MBZUAI), Daria Galimzianova (MBZUAI), Nikolai Rozanov (MBZUAI, Imperial College London… ·

    ThinkBooster:LLM推理的无缝测试时尺度统一框架

    arXiv:2606.06915v1 Announce Type: cross Abstract: Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute during inference, e.g., via multi-sample generation and verifier-based rerankin…

  143. arXiv cs.AI TIER_1 English(EN) · Debjyoti Saha Roy, Byron C. Wallace, Javed A. Aslam ·

    表征再提炼:大输出空间中的机制推理

    arXiv:2606.06840v1 Announce Type: cross Abstract: Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevant options from hundreds of thousands to millions of candidate labels. We investi…

  144. arXiv cs.AI TIER_1 English(EN) · Tengyao Tu, Yulin Li, Hui-Ling Zhen, Libo Qin, Zhoujun Wei, Jinghua Piao, Zhuotao Tian, Yong Li, Min Zhang ·

    DyCon:通过演进难度建模实现动态推理控制

    arXiv:2606.07108v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as …

  145. arXiv cs.CL TIER_1 English(EN) · Donald Ye, Max Loffgren, Om Kotadia, Linus Wong, Jonas Rohweder ·

    Chain-of-Thought推理中忠实度衰减的机制证据

    arXiv:2602.11201v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) explanations are widely used to interpret how language models solve complex problems, yet it remains unclear whether these step-by-step explanations reflect how the model actually reaches its answer, or me…

  146. arXiv cs.CL TIER_1 English(EN) · Xinze Li, Yuqing Lan, Zhenghao Liu, Haidong Xin, Yukun Yan, Shuo Wang, Zheni Zeng, Sen Mei, Ge Yu, Maosong Sun ·

    SEEK:通过内部推理草图引导 RAG 的 LLM 推理

    arXiv:2601.09402v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge into the generation process. Benefiting from the reasoning capabilities of LLMs, existing methods have leveraged such…

  147. arXiv cs.CL TIER_1 English(EN) · Yongliang Miao, Fengyuan Liu, Wei Shi, Yanguang Liu, Fei Sun, Na Zou, Mengnan Du ·

    RASFT:面向推理的滚动自适应监督微调

    arXiv:2606.07006v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations, often treating a single expert trajectory as the target behavior. However, reaso…

  148. arXiv cs.CL TIER_1 English(EN) · Zhixuan He, Yue Feng ·

    何时深入思考:LLM推理的抑制性审议

    arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is computationally expensive and often unnecessary. We propose IDPR, a framework for…

  149. arXiv cs.AI TIER_1 English(EN) · Raman Saparkhan, Majd Hawasly, Md Rizwan Parvez, Mohammad Raza ·

    仅需两个样本即可实现自洽性:CoT-PoT重排提升LLM推理效率

    arXiv:2604.17433v2 Announce Type: replace-cross Abstract: Self-consistency (SC) is a popular technique for improving the reasoning accuracy of large language models by aggregating multiple sampled outputs, but it comes at a high computational cost due to extensive sampling. We in…

  150. arXiv cs.AI TIER_1 English(EN) · Yuxiang Chen, Jun Wang ·

    人类与 DeepSeek-R1 LLM 数学推理的全面解剖

    arXiv:2606.07410v1 Announce Type: cross Abstract: The emergence of "Aha moments" in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems genuinely reason or merely imitate the appearance of reasoning. We conduct a comprehensive em…

  151. arXiv cs.AI TIER_1 English(EN) · Rahul Nair, Chun Tao ·

    微调陷阱:评估负迁移和 PEFT 在低于 10 亿参数数学推理中的作用

    arXiv:2606.06920v1 Announce Type: cross Abstract: Deploying Small Language Models (SLMs) on edge devices requires efficient fine-tuning strategies that adapt models to new tasks without degrading their general capabilities. In this study, we benchmark five sub-1B models (135M-1B)…

  152. arXiv cs.AI TIER_1 English(EN) · Hejun Wu ·

    将大型语言模型推理提炼为可解释的策略树,以实现人机协作

    Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of work. Most prior work relies on multi-agent reinforcement learning (MARL) to learn black-box policies, which limits interpretabil…

  153. arXiv cs.AI TIER_1 English(EN) · Subramanyam Sahoo ·

    用于诊断推理模型中未知未知信息的结构化无知证书的校准

    Large language models frequently fail in a characteristic way: rather than acknowledging ignorance, they produce fluent but incorrect answers to questions that lie beyond their knowledge boundaries. We introduce \textbf{Structured Ignorance Certificates} (SICs), a JSON-formatted …

  154. arXiv cs.CL TIER_1 English(EN) · Xueru Zhang ·

    TLRD:教大型语言模型通过三级原理提炼来推理表格数据

    Tabular data is a primary medium for storing real-world information, driving many industrial applications of machine learning. Traditional predictors achieve strong predictive performance but do not provide readable, case-specific explanations essential for decision-making. Large…

  155. arXiv cs.AI TIER_1 English(EN) · Xiaopeng Yuan, Haibo Jin, Ye Yu, Peng Kuang, Lijun Yu, Yushun Dong, Haohan Wang ·

    通过测试时重构实现潜在推理的闭环

    arXiv:2606.06252v1 Announce Type: new Abstract: Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a discrete communication bottleneck. However, this shift also removes a key advantag…

  156. arXiv cs.AI TIER_1 English(EN) · Jiate Liu, Zebin Chen, Shaobo Qiao, Mingchen Ju, Danting Zhang, Bocheng Han, Shuyue Yu, Xin Shu, Jinglin Wu, Dong Wen, Xin Cao, Guanfeng Liu, Zhengyi Yang ·

    A2RAG:自适应代理图检索,用于成本感知和可靠的推理

    arXiv:2601.21162v2 Announce Type: replace-cross Abstract: Graph Retrieval-Augmented Generation (Graph-RAG) enhances multihop question answering by organizing corpora into knowledge graphs and routing evidence through relational structure. However, practical deployments face two p…

  157. arXiv cs.AI TIER_1 English(EN) · Hamed Nejat, Alexander Maier, Jesse Spencer-Smith, Andr\'e M. Bastos ·

    本体约束的多大型语言模型对预测性处理文献中假设支持的评分

    arXiv:2606.05206v1 Announce Type: cross Abstract: Fragmentation is common in interdisciplinary fields with diverse methods and theoretical commitments. Predictive coding neuroscience is a clear example: its literature spans computational theory, electrophysiology, imaging, behavi…

  158. arXiv cs.LG TIER_1 English(EN) · Jun Wang ·

    人类与 DeepSeek-R1 LLM 数学推理的全面解剖

    The emergence of "Aha moments" in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems genuinely reason or merely imitate the appearance of reasoning. We conduct a comprehensive empirical comparison between model and human reasoni…

  159. arXiv cs.AI TIER_1 English(EN) · Min Zhang ·

    DyCon:通过演进式难度建模实现动态推理控制

    Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as "overthinking". Existing methods to mitigate thi…

  160. arXiv cs.CL TIER_1 English(EN) · Mengnan Du ·

    RASFT:面向推理的滚动自适应监督微调

    Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations, often treating a single expert trajectory as the target behavior. However, reasoning is not simple path imitation: rigidly followi…

  161. arXiv cs.AI TIER_1 English(EN) · Chun Tao ·

    微调陷阱:评估负迁移和 PEFT 在低于 10 亿参数数学推理中的作用

    Deploying Small Language Models (SLMs) on edge devices requires efficient fine-tuning strategies that adapt models to new tasks without degrading their general capabilities. In this study, we benchmark five sub-1B models (135M-1B) on mathematical reasoning tasks and uncover a cri…

  162. arXiv cs.CL TIER_1 English(EN) · Artem Shelmanov ·

    ThinkBooster:LLM推理的无缝测试时缩放的统一框架

    Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute during inference, e.g., via multi-sample generation and verifier-based reranking. Existing TTC scaling strategies and reasoning s…

  163. arXiv cs.LG TIER_1 English(EN) · Nirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor, Roi Reichart ·

    使用反事实链和因果图实现大语言模型可解释性

    arXiv:2606.05972v1 Announce Type: new Abstract: Causal graphs provide a high-level language for making mechanisms transparent. Recent work uses Large Language Models (LLMs) to recover causal graphs of external-world processes. Instead, in this paper, we use causal graphs to model…

  164. arXiv cs.LG TIER_1 English(EN) · Locke Cai, Max Ryabinin, Ivan Provilkov ·

    逃离验证器:通过演示学习推理

    arXiv:2511.21667v4 Announce Type: replace Abstract: Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-world reasoning-intensive tasks lack verifiers, despite offering abundant expert demons…

  165. arXiv cs.LG TIER_1 English(EN) · Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp Hochreiter ·

    RREDCoT:面向推理模型的片段级奖励再分配

    arXiv:2606.06475v1 Announce Type: new Abstract: Recent advancements in reasoning language models have been driven by Reinforcement Learning (RL) fine-tuning. Most often, these rely on the Group Relative Policy Optimization (GRPO) algorithm or modifications thereof to steer the mo…

  166. arXiv cs.LG TIER_1 English(EN) · Rohan Siva, Neel P. Bhatt, Yunhao Yang, Seoyoung Lee, Nishant Gadde, Christian Ellis, Alvaro Velasquez, Zhangyang Wang, Ufuk Topcu ·

    关注物体赋能而非物体本身:用于可供性推理的功能性潜在空间

    arXiv:2606.05533v1 Announce Type: new Abstract: Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object appearances (e.g., recognizing a "cart" based on how it looks). However, planning r…

  167. arXiv cs.CL TIER_1 English(EN) · Ashima Suvarna, Kendrick Phan, Mehrab Beikzadeh, Hritik Bansal, Saadia Gabriel ·

    SUPERNOVA:利用自然指令上的强化学习在大型语言模型中引发通用推理

    arXiv:2604.08477v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has substantially improved reasoning in formal domains such as mathematics and code, but extending these gains beyond STEM remains challenging. Extending RLVR beyond ST…

  168. arXiv cs.CL TIER_1 English(EN) · Chengwei Wei, Jung-jae Kim, Longyin Zhang, Shengkai Chen, Nancy F. Chen ·

    InfoDensity:奖励信息密集型轨迹以实现高效推理

    arXiv:2603.17310v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary computational cost. While existing reinforcement learning approaches address th…

  169. arXiv cs.CL TIER_1 English(EN) · Zhenyuan Guo, Tong Chen, Wenlong Meng, Chen Gong, Xin Yu, Chengkun Wei, Wenzhi Chen ·

    面向大型推理模型的高效推理动态思维-Token选择

    arXiv:2601.18383v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. However, these extended generations incur substantial memory footprint and comput…

  170. arXiv cs.CL TIER_1 English(EN) · Ayoung Lee, Ryan Sungmo Kwon, Peter Railton, Lu Wang ·

    CLASH:从多角度评估语言模型在判断高风险困境中的表现

    arXiv:2504.10823v4 Announce Type: replace Abstract: Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been limited to everyday scenarios. To close this gap, we introduce CLASH (Character pe…

  171. arXiv cs.CL TIER_1 English(EN) · Maxime Griot, Paul Steven Scotti, Tanishq Mathew Abraham ·

    Compress-Distill:用于高效知识蒸馏的推理轨迹压缩

    arXiv:2606.05988v1 Announce Type: cross Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression of such traces before knowledge distillation. Two teachers, Qwen3.5-397B-A17B an…

  172. arXiv cs.CL TIER_1 English(EN) · Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu ·

    使用归一化流的潜在推理

    arXiv:2606.06447v1 Announce Type: new Abstract: Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and comm…

  173. arXiv cs.CL TIER_1 English(EN) · Jinyang Zhang, Hongxin Ding, Yue Fang, Weibin Liao, Muyang Ye, Junfeng Zhao, Yasha Wang ·

    The Tell-Tale Norm: $\ell_2$ Magnitude as a Signal for Reasoning Dynamics in Large Language Models

    arXiv:2606.06188v1 Announce Type: new Abstract: Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reasoning dynamics remains underexplored. We bridge this gap by demonstrating that th…

  174. arXiv cs.CL TIER_1 English(EN) · Liting Zhang, Shiwan Zhao, Xuyang Zhao, Zichen Xu, Jianye Wang, Qicheng Li ·

    TARPO:通过动作路由策略优化实现令牌级隐式推理

    arXiv:2606.05859v1 Announce Type: new Abstract: Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive reasoning by operating over continuous representations. However, the inherently det…

  175. arXiv cs.CL TIER_1 English(EN) · Jinu Lee, Shivam Agarwal, Amruta Parulekar, Siddarth Madala, Dilek Hakkani-Tur, Julia Hockenmaier ·

    ReasoningFlow: 理解LLM推理轨迹的论述结构

    arXiv:2606.05402v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process. We introduce ReasoningFlow, a frame…

  176. arXiv cs.CL TIER_1 English(EN) · Ryan Solgi, Jiayi Tian, Zheng Zhang ·

    LoRi:低秩蒸馏用于隐式推理

    arXiv:2606.05315v1 Announce Type: new Abstract: Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting. We empirically find that hidden-state reasoning trajectories exhibit low-rank structure. …

  177. arXiv cs.CL TIER_1 English(EN) · Javed A. Aslam ·

    表征然后提炼:大输出空间中的机制推理

    Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevant options from hundreds of thousands to millions of candidate labels. We investigate how they achieve this mechanistically. We cha…

  178. arXiv cs.CL TIER_1 English(EN) · Yue Feng ·

    何时深入思考:LLM推理的抑制性审议

    Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is computationally expensive and often unnecessary. We propose IDPR, a framework for response-conditioned inhibitory deliberation. I…

  179. arXiv cs.CL TIER_1 English(EN) · Mykel J. Kochenderfer ·

    语言模型如何失效:已承诺和持续推理失败的 Token 级签名

    Failures in language model reasoning emerge through distinct processes that leave identifiable signatures in the reasoning trace. We characterize these failures using token-level uncertainty signals, finding they arise through two empirically distinguishable processes. The first …

  180. arXiv cs.AI TIER_1 English(EN) · Sepp Hochreiter ·

    RREDCoT:用于推理模型的片段级奖励再分配

    Recent advancements in reasoning language models have been driven by Reinforcement Learning (RL) fine-tuning. Most often, these rely on the Group Relative Policy Optimization (GRPO) algorithm or modifications thereof to steer the models to produce Chain-of-Thought (CoT) traces. T…

  181. arXiv cs.CL TIER_1 English(EN) · Jiatao Gu ·

    使用归一化流的潜在推理

    Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning …

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用归一化流的潜在推理

    Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning …

  183. arXiv cs.AI TIER_1 English(EN) · Haohan Wang ·

    通过测试时重构实现潜在推理的闭环

    Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a discrete communication bottleneck. However, this shift also removes a key advantage of textual reasoning: intermediate states are …

  184. arXiv cs.CL TIER_1 English(EN) · Yasha Wang ·

    The Tell-Tale Norm: $\ell_2$ Magnitude as a Signal for Reasoning Dynamics in Large Language Models

    Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reasoning dynamics remains underexplored. We bridge this gap by demonstrating that the l2 norm of hidden states serves as an endogeno…

  185. arXiv cs.CL TIER_1 English(EN) · Tanishq Mathew Abraham ·

    Compress-Distill:用于高效知识蒸馏的推理轨迹压缩

    Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression of such traces before knowledge distillation. Two teachers, Qwen3.5-397B-A17B and gpt-oss-120B, generate about 283k correct traces…

  186. arXiv cs.CL TIER_1 English(EN) · Qicheng Li ·

    TARPO:通过动作路由策略优化实现令牌级潜在显式推理

    Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive reasoning by operating over continuous representations. However, the inherently deterministic nature of continuous representations …

  187. arXiv cs.AI TIER_1 (AF) · Anshul Nayak, Shahil Shaik, Yue Wang ·

    面向类人推理的信念感知VLM模型

    arXiv:2604.09686v2 Announce Type: replace Abstract: Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic environments. Recent advances in Vision Language Models (VLMs) and Vision Langu…

  188. arXiv cs.AI TIER_1 English(EN) · Wang Yang, Xiang Yue, Vipin Chaudhary, Xiaotian Han ·

    推测性思维:在推理时使用大型模型指导增强小型模型的推理能力

    arXiv:2504.12329v2 Announce Type: replace-cross Abstract: Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still suffers from inefficient, overly lengthy outputs. We introduce Speculative Thinkin…

  189. arXiv cs.AI TIER_1 English(EN) · Zheng Du, Hao Kang, Song Han, Tushar Krishna, Ligeng Zhu ·

    OckBench:衡量大型语言模型推理效率

    arXiv:2511.05722v3 Announce Type: replace-cross Abstract: Large language models (LLMs) such as GPT-5 and Gemini 3 have pushed the frontier of automated reasoning and code generation. Yet current benchmarks emphasize accuracy and output quality, neglecting a critical dimension: ef…

  190. arXiv cs.AI TIER_1 English(EN) · Wang Yang, Debargha Ganguly, Xinpeng Li, Chaoda Song, Shouren Wang, Vikash Singh, Vipin Chaudhary, Xiaotian Han ·

    Mid-Think:通过令牌级触发器进行无训练的中间预算推理

    arXiv:2601.07036v2 Announce Type: replace-cross Abstract: Hybrid reasoning language models are commonly controlled through high-level Think/No-think instructions to regulate reasoning behavior, yet we found that such mode switching is largely driven by a small set of trigger toke…

  191. arXiv cs.AI TIER_1 English(EN) · Ethan Mendes, Jungsoo Park, Alan Ritter ·

    让专家推理可学习:通过自我蒸馏

    arXiv:2602.02405v2 Announce Type: replace-cross Abstract: Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct solution to be reinforced or the existence of a stronger model able to solve the probl…

  192. arXiv cs.AI TIER_1 English(EN) · Jonas Petersen, Camilla Mazzoleni, Gian-Alessandro Lombardi, Federico Martelli, Riccardo Maggioni ·

    什么结构归纳偏倚有助于 Transformer 推理知识图谱?一项使用 Tabula RASA 的研究

    arXiv:2602.02834v4 Announce Type: replace-cross Abstract: What structural inductive bias helps transformers reason over knowledge graphs? Through controlled ablations of a minimal transformer modification with four independently removable components (sparse adjacency masking, edg…

  193. arXiv cs.CL TIER_1 English(EN) · Chongyang He, Rui Zhang, Zixuan Wang, Xin Li ·

    学习如何学习:用于小型语言模型推理的SFT-then-RL阶段特定数据集

    arXiv:2606.04466v1 Announce Type: new Abstract: Post-training Small Language Models (SLMs) for reasoning typically follows an SFT-then-RL pipeline, yet existing work rarely considers what data should be learned at each stage. We argue that data strategy should be aligned with the…

  194. arXiv cs.CL TIER_1 English(EN) · Haoran Zhang, Yafu Li, Zhi Wang, Zhilin Wang, Shunkai Zhang, Xiaoye Qu, Yu Cheng ·

    复杂推理的表征、评估与优化

    arXiv:2602.08498v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) increasingly rely on reasoning traces with complex internal structures. However, existing work lacks a unified answer to three fundamental questions: (1) what defines high-quality reasoning, (2) how…

  195. arXiv cs.CL TIER_1 English(EN) · Siqi Fan, Minghao Li, Xiaoqian Ma, Xiusheng Huang, Zhuo Chen, Bowen Qin, Liujie Zhang, Shuo Shang, Weihang Chen ·

    Hint Tuning:数据越少,推理能力越强

    arXiv:2605.08665v2 Announce Type: replace Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning uniformly regardless of problem difficulty. We propose Hint Tuning, a data-e…

  196. arXiv cs.CL TIER_1 English(EN) · Sanket Badhe, Deep Shah ·

    Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning

    arXiv:2602.21103v2 Announce Type: replace Abstract: Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency and substantial test-time inference costs. The standard alternative, fine-tuning smaller models, often sacrifices…

  197. arXiv cs.CL TIER_1 English(EN) · Yelysei Bondarenko, Thomas Hehn, Rob Hesselink, Romain Lepert, Fabio Valerio Massoli, Evgeny Mironov, Leyla Mirvakhabova, Tribhuvanesh Orekondy, Spyridon Stasis, Andrey Kuzmin, Anna Kuzina, Markus Nagel, Ankita Nayak, Corrado Rainone, Ork de Rooij, Paul … ·

    边缘高效推理

    arXiv:2603.16867v2 Announce Type: replace-cross Abstract: Large language models (LLMs) with chain-of-thought reasoning achieve state-of-the-art performance across complex problem-solving tasks, but their verbose reasoning traces and large context requirements make them impractica…

  198. arXiv cs.LG TIER_1 English(EN) · Tiehua Mei, Minxuan Lv, Leiyu Pan, Zhenpeng Su, Hongru Hou, Hengrui Chen, Ao Xu, Deqing Yang ·

    良好的推理造就良好的演示:通过上下文强化学习进行隐式推理质量监督

    arXiv:2603.09803v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in large language models but treats all correct solutions equally, potentially reinforcing flawed traces that arrive at correct answers by chance. We obser…

  199. arXiv cs.LG TIER_1 English(EN) · Gleb Rodionov, Roman Garipov, George Yakushev ·

    推理转变:上下文如何悄然缩短大语言模型的推理能力

    arXiv:2604.01161v2 Announce Type: replace Abstract: Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness…

  200. arXiv cs.AI TIER_1 English(EN) · Jingbo Wen, Liang He, Ziqi He ·

    并非所有错误都均等:考虑后果的推理计算分配

    arXiv:2606.04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks. Existing methods generally drive this allocation by predicted difficulty a…

  201. arXiv cs.AI TIER_1 English(EN) · Yuhan Yang, Ruipu Li, Alexander Rodr\'iguez ·

    模拟、推理、决策:利用大型语言模型进行模拟驱动决策的科学推理

    arXiv:2606.04505v1 Announce Type: new Abstract: Scientific simulators are increasingly being integrated into LLM-driven systems for high-stakes simulation-driven decision-making. However, existing frameworks primarily use LLMs to generate, calibrate, or execute simulators, treati…

  202. arXiv cs.AI TIER_1 English(EN) · Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi ·

    FALSIFYBENCH:使用规则发现游戏评估大型语言模型的归纳推理能力

    arXiv:2606.04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in forms of inductive reasoning relevant to scientific discovery remains an open quest…

  203. arXiv cs.AI TIER_1 English(EN) · Guangyao Dou, William Jurayj, Nils Holzenberger, Benjamin Van Durme ·

    DAR:使用智能体工具进行义务推理

    arXiv:2606.05009v1 Announce Type: cross Abstract: Deontic reasoning is the task of answering questions by applying explicit rules and policies to case-specific facts, for example computing tax liability under a statute or determining the outcome of an immigration appeal. A key te…

  204. arXiv cs.AI TIER_1 English(EN) · Zehua Cheng, Wei Dai, Jiahao Sun ·

    用于鲁棒推理蒸馏的不变梯度对齐

    arXiv:2606.05025v1 Announce Type: cross Abstract: Large language models (LLMs) suffer from shortcut learning: they systematically fail on out-of-distribution (OOD) inputs whose semantic surface differs from training data, even when the logical structure is identical. This undermi…

  205. arXiv cs.AI TIER_1 English(EN) · Wang Yang, Zirui Liu, Hongye Jin, Qingyu Yin, Vipin Chaudhary, Xiaotian Han ·

    更长上下文,更深思考:揭示长上下文能力在推理中的作用

    arXiv:2505.17315v2 Announce Type: replace Abstract: Recent language models exhibit strong reasoning capabilities, yet the influence of long-context capacity on reasoning remains underexplored. In this work, we hypothesize that current limitations in reasoning stem, in part, from …

  206. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用归一化流的潜在推理

    Latent reasoning framework using normalizing flows preserves autoregressive generation advantages while enabling efficient, probabilistic intermediate computation in large language models.

  207. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用反事实链和因果图实现大语言模型可解释性

    Causal graphs are used to model large language model inference processes, enabling transparent visualization of how models perceive and organize high-level concepts for predictions through a four-phase method involving concept discovery, mapping, and MCMC-inspired counterfactual …

  208. Hugging Face Daily Papers TIER_1 English(EN) ·

    Compress-Distill:用于高效知识蒸馏的推理轨迹压缩

    Post-hoc compression of reasoning traces reduces computational costs and inference lengths while maintaining high accuracy, offering an accuracy-efficiency trade-off in knowledge distillation.

  209. arXiv cs.LG TIER_1 English(EN) · Jiahao Sun ·

    Invariant Gradient Alignment for Robust Reasoning Distillation

    Large language models (LLMs) suffer from shortcut learning: they systematically fail on out-of-distribution (OOD) inputs whose semantic surface differs from training data, even when the logical structure is identical. This undermines knowledge distillation pipelines that transfer…

  210. arXiv cs.CL TIER_1 English(EN) · Benjamin Van Durme ·

    DAR:基于Agentic Harnesses的义务推理

    Deontic reasoning is the task of answering questions by applying explicit rules and policies to case-specific facts, for example computing tax liability under a statute or determining the outcome of an immigration appeal. A key technical challenge for LLM-based deontic reasoning …

  211. arXiv cs.AI TIER_1 English(EN) · Raffaella Bernardi ·

    FALSIFYBENCH:使用规则发现游戏评估LLM的归纳推理能力

    Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in forms of inductive reasoning relevant to scientific discovery remains an open question. In this work, we introduce FALSIFYBENCH, an…

  212. Hugging Face Daily Papers TIER_1 English(EN) ·

    学习如何学习:用于小型语言模型推理的SFT-then-RL分阶段数据集

    Post-training Small Language Models (SLMs) for reasoning typically follows an SFT-then-RL pipeline, yet existing work rarely considers what data should be learned at each stage. We argue that data strategy should be aligned with the distinct roles of SFT and RL: SFT is better sui…

  213. arXiv cs.CL TIER_1 English(EN) · Xin Li ·

    学习如何学习:用于小型语言模型推理的SFT-then-RL分阶段数据集

    Post-training Small Language Models (SLMs) for reasoning typically follows an SFT-then-RL pipeline, yet existing work rarely considers what data should be learned at each stage. We argue that data strategy should be aligned with the distinct roles of SFT and RL: SFT is better sui…

  214. arXiv cs.AI TIER_1 English(EN) · Cl\'ement Yvernes, Emilie Devijver, Marianne Clausel, Eric Gaussier ·

    通过推导图揭示 Do-Calculus 推理的结构

    arXiv:2606.03719v1 Announce Type: new Abstract: The do-calculus defines a general system of inference for interventional queries, allowing causal quantities to be transformed through successive applications of its rules. This process induces a rich space of equivalent interventio…

  215. arXiv cs.AI TIER_1 English(EN) · Zhengyi Zhao, Shubo Zhang, Huimin Wang, Zezhong Wang, Yutian Zhao, Yefeng Zheng, Binyang Li, Yulan He, Kam-Fai Wong, Xian Wu ·

    利用辅助约束解决大型推理模型的指令遵循问题

    arXiv:2606.03624v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions, either by failing to satisfy individual constraints or by struggling to balance …

  216. arXiv cs.AI TIER_1 English(EN) · Chuang Yu, Jinmiao Zhao, Mingxuan Zhao, Yunpeng Liu, Xiujun Shu, Yuanhao Feng, Bo Wang, Xiangyu Yue ·

    MIND:面向多模态大模型的集成判别式多推理框架

    arXiv:2512.05530v2 Announce Type: replace Abstract: Recently, multimodal large language models (MLLMs) have been widely applied to reasoning tasks. However, they suffer from limited multi-rationale semantic modeling, insufficient logical robustness, and susceptibility to misleadi…

  217. arXiv cs.AI TIER_1 English(EN) · Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu, Arman Cohan ·

    量化大型推理模型中忠实置信度表达

    arXiv:2606.03969v1 Announce Type: cross Abstract: Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode. This ch…

  218. arXiv cs.AI TIER_1 English(EN) · Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley ·

    Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

    arXiv:2606.03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking l…

  219. arXiv cs.AI TIER_1 English(EN) · Dongwon Jung, Peng Shi, Yi Zhang, Junshan Zhang, Muhao Chen ·

    自适应潜在代理推理

    arXiv:2606.02871v1 Announce Type: cross Abstract: Large reasoning models improve performance by generating extended chain-of-thought (CoT) reasoning, but this behavior becomes inefficient when applied to LLM agents. Current LLM agents often generate verbose textual reasoning at e…

  220. arXiv cs.AI TIER_1 English(EN) · Eric Cho, Shawn Huang, Alice Lu, Andy Lyu ·

    Hedge-Bench:在涉及金融推理的困难、现实任务上对代理进行基准测试

    arXiv:2606.03918v1 Announce Type: new Abstract: AI agents can increasingly handle the mechanical tasks of financial analysis: retrieving documents, calculating formulas, updating spreadsheets. The harder, more valuable challenge is reasoning through the open-ended questions that …

  221. arXiv cs.AI TIER_1 English(EN) · Ayushi Chadha ·

    何时重新规划:分层潜在推理中的子目标持久性

    arXiv:2606.03741v1 Announce Type: new Abstract: Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never coheres into multi-step structure; commit too long and the plan goes stale. We study this st…

  222. arXiv cs.AI TIER_1 English(EN) · Hongyu Guo, Hao Li, He Cao, Gongbo Zhang, Li Yuan ·

    从答案到状态:大语言模型化学推理的可验证过程级评估

    arXiv:2606.03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while…

  223. arXiv cs.AI TIER_1 English(EN) · Simone Caldarella, Davide Talon, Rahaf Aljundi, Elisa Ricci, Massimiliano Mancini ·

    超越答案的思考:评估大型推理模型中的有害过度思考

    arXiv:2606.02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assumption that longer reasoning is consistently beneficial remains under-examined. …

  224. arXiv cs.AI TIER_1 English(EN) · Zhihan Lei, Jiarui Yan, Joshua Momo, William W. Cohen ·

    从智能体轨迹中诱导推理原语

    arXiv:2606.02994v1 Announce Type: new Abstract: ReAct-style LLM agents often rediscover the same reasoning routines across problems, yet leave those routines trapped in transient scratchpads. We introduce Reasoning Primitive Induction, a single-pass method that mines successful R…

  225. arXiv cs.AI TIER_1 English(EN) · Ziyan Liu, Xueda Shen, Yuzhe Gu, Songyang Gao, Kuikun Liu, Guangran Cheng, Chengqi Lyu, Dahua Lin, Wenwei Zhang, Kai Chen ·

    ThoughtFold:通过内省偏好学习折叠推理链

    arXiv:2606.03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (CoTs). However, since long CoTs naturally contain trial and errors and mainstream …

  226. arXiv cs.LG TIER_1 English(EN) · Aijia Cheng, Kailong Wang, Ling Shi, Yongxin Zhao ·

    R2IF:通过复合奖励实现可解释大模型函数调用的推理与决策对齐

    arXiv:2604.20316v2 Announce Type: replace Abstract: Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment between reasoning processes and tool-call decisions. We propose R2IF, a reasoning…

  227. arXiv cs.LG TIER_1 English(EN) · Ziyue Wang, Aomufei Yuan, Yongfu Zhu, Shuai Dong, Wenpu Liu, Yiran Yao, Weichu Xie, Yuqi Xu, Caoyuan Ma, Wenqi Shao, Xiaoying Zhang, Nan Duan, Jiaqi Wang ·

    正义即力量:对齐已验证的隐藏状态可增强强化学习推理能力

    arXiv:2606.03234v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has become the dominant approach for improving mathematical reasoning in large language models, yet current methods reduce each correct rollout to a single reward bit, ignoring t…

  228. arXiv cs.CL TIER_1 English(EN) · Jiaxi Bi, Tongxu Luo, Wenyu Du, Zhengyang Tang, Benyou Wang ·

    及时止损!学习早期剪枝以实现高效并行推理

    arXiv:2604.16029v2 Announce Type: replace Abstract: Parallel reasoning enhances Large Reasoning Models (LRMs) but incurs prohibitive costs due to futile paths caused by early errors. To mitigate this, path pruning at the prefix level is essential, yet existing research remains fr…

  229. arXiv cs.CL TIER_1 English(EN) · Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen ·

    世界模型与语言模型相遇:论具体推理与抽象推理的互补性

    arXiv:2606.03603v1 Announce Type: cross Abstract: World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can generate concrete visual rollouts of possible futures, w…

  230. arXiv cs.AI TIER_1 English(EN) · Dani Roytburg, Shreya Sridhar, Daphne Ippolito ·

    衡量推理模型弱到强可读性

    arXiv:2603.20508v2 Announce Type: replace-cross Abstract: Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such as inter-model monitoring or distillation into smaller models. When agents at d…

  231. arXiv cs.AI TIER_1 English(EN) · Xinwu Ye, Yicheng Mao, Yuxuan Liao, Jia Zhang, Yimeng Liu, Li Hao, Fang Wu, Zhiwei Li, Zehong Wang, Zhiyuan Liu, Zhenfei Yin, Li Yuan, Philip Torr, Huan Sun, xiangxiang Zeng, Mengdi Wang, Le Cong, Shenghua Gao, Xiangru Tang ·

    LatentChem:从文本CoT到化学推理中的潜在思维

    arXiv:2602.07075v5 Announce Type: replace-cross Abstract: Current chemical large language models (LLMs) predominantly rely on explicit Chain-of-Thought (CoT) to solve complex reasoning problems. However, forcing nonverbal tacit chemical logic into discrete natural language impose…

  232. arXiv cs.AI TIER_1 English(EN) · Yuchen Yan, Liang Jiang, Jin Jiang, Shuaicheng Li, Zujie Wen, Zhiqiang Zhang, Jun Zhou, Jian Shao, Yueting Zhuang, Yongliang Shen ·

    InftyThink+: 通过强化学习实现有效且高效的无限视野推理

    arXiv:2602.06960v3 Announce Type: replace-cross Abstract: Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, context length limits, and degraded reasoning due to lost-in-the-middle effects. …

  233. Hugging Face Daily Papers TIER_1 English(EN) ·

    DAR:使用代理工具进行义务推理

    Deontic reasoning tasks require applying complex rules and policies, and an agentic approach enables models to dynamically access statutes, showing mixed performance improvements across different model strengths.

  234. arXiv cs.AI TIER_1 English(EN) · Arman Cohan ·

    量化大型推理模型中忠实置信度的表达

    Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode. This challenge is key for large reasoning models (LRMs), …

  235. arXiv cs.AI TIER_1 English(EN) · Julian McAuley ·

    Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

    Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by shortening, early-stopping, or compressin…

  236. arXiv cs.AI TIER_1 English(EN) · Andy Lyu ·

    Hedge-Bench:在涉及金融推理的困难、现实任务上对代理进行基准测试

    AI agents can increasingly handle the mechanical tasks of financial analysis: retrieving documents, calculating formulas, updating spreadsheets. The harder, more valuable challenge is reasoning through the open-ended questions that define expert Analyst work. Existing benchmarks …

  237. arXiv cs.AI TIER_1 English(EN) · Ayushi Chadha ·

    何时重新规划:分层潜在推理中的子目标持久性

    Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never coheres into multi-step structure; commit too long and the plan goes stale. We study this stability-adaptivity tradeoff in the latent reason…

  238. arXiv cs.AI TIER_1 English(EN) · Eric Gaussier ·

    通过推导图揭示 Do-Calculus 推理的结构

    The do-calculus defines a general system of inference for interventional queries, allowing causal quantities to be transformed through successive applications of its rules. This process induces a rich space of equivalent interventional expressions, but combining and ordering thes…

  239. arXiv cs.AI TIER_1 English(EN) · Li Yuan ·

    从答案到状态:大型语言模型化学推理的可验证过程级评估

    Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing…

  240. arXiv cs.AI TIER_1 English(EN) · Xian Wu ·

    利用辅助约束解决大型推理模型的指令遵循问题

    Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions, either by failing to satisfy individual constraints or by struggling to balance competing constraints simultaneously. We formali…

  241. Hugging Face Daily Papers TIER_1 English(EN) ·

    利用辅助约束解决大型推理模型的指令遵循问题

    Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions, either by failing to satisfy individual constraints or by struggling to balance competing constraints simultaneously. We formali…

  242. Hugging Face Daily Papers TIER_1 English(EN) ·

    世界模型与语言模型相遇:论具体推理与抽象推理的互补性

    Controlled concrete reasoning combines visual simulation with abstract reasoning through a training method that uses privileged future information to improve prediction accuracy and robustness.

  243. Hugging Face Daily Papers TIER_1 English(EN) ·

    世界模型与语言模型相遇:论具体推理与抽象推理的互补性

    World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can generate concrete visual rollouts of possible futures, while MLLMs can reason abstractly over questions, g…

  244. arXiv cs.CL TIER_1 English(EN) · Jianbing Shen ·

    世界模型与语言模型相遇:论具体推理与抽象推理的互补性

    World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can generate concrete visual rollouts of possible futures, while MLLMs can reason abstractly over questions, g…

  245. Hugging Face Daily Papers TIER_1 English(EN) ·

    ThoughtFold:通过内省偏好学习折叠推理链

    ThoughtFold addresses over-thinking in large reasoning models by using fine-grained preference learning to identify and eliminate redundant explorations in chain-of-thought reasoning processes.

  246. arXiv cs.CL TIER_1 English(EN) · Shuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma, Qingyang Liu, Zhaohe Liao, Yibo Miao, Li Niu ·

    解锁潜在推理的黑箱:一种以可解释性为导向的干预方法

    arXiv:2606.01243v1 Announce Type: new Abstract: Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over explicit Chain-of-Thought (CoT). However, the opacity of these continuous thought …

  247. arXiv cs.LG TIER_1 English(EN) · Xuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu, Zhenya Huang, Ning Miao ·

    用于推理过程解释的步进稀疏自编码器

    arXiv:2603.03031v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved strong complex reasoning capabilities through Chain-of-Thought (CoT) reasoning. However, their reasoning patterns remain too complicated to analyze. While Sparse Autoencoders (SAEs) hav…

  248. arXiv cs.LG TIER_1 English(EN) · Sanae Lotfi, Polina Kirichenko, Steven Li, Zechun Liu ·

    量化推理模型认为它们需要更长时间思考,但事实并非如此

    arXiv:2606.00206v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ reduces accurac…

  249. arXiv cs.LG TIER_1 English(EN) · Arif Hassan Zidan, Yi Pan, Hanqi Jiang, Ruiyu Yan, Wei Ruan, Zihao Wu, Lifeng Chen, Weihang You, Xinliang Li, Bowen Chen, Huawen Hu, Peilong Wang, Sizhuang Liu, Jing Zhang, Siyuan Li, Zhengliang Liu, Yu Bao, Lin Zhao, Lichao Sun, Dajiang Zhu, Xiang Li, J… ·

    世界模型:架构、方法论、推理范式和应用的综合调查

    arXiv:2606.00133v1 Announce Type: new Abstract: World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artificial general intelligence, enabling agents to predict, plan, and reason within lea…

  250. arXiv cs.CL TIER_1 English(EN) · Kasidit Sermsri, Teerapong Panboonyuen ·

    GateKD: 置信度门控闭环蒸馏实现鲁棒推理

    arXiv:2605.13136v2 Announce Type: replace Abstract: Distilling multi-step reasoning abilities from large language models (LLMs) into compact student models remains challenging due to noisy rationales, hallucinated supervision, and static teacher-student interactions. Existing rea…

  251. arXiv cs.CL TIER_1 English(EN) · Songze Li, Zhiqiang Liu, Zhaoyan Gong, Xiaoke Guo, Zhongpu Bo, Zhengke Gui, Lei Liang, Huajun Chen, Wen Zhang ·

    Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning

    arXiv:2511.07910v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve excellent performance in natural language reasoning tasks through pre-training on vast unstructured text, enabling them to understand the logic in natural language and generate logic-consiste…

  252. arXiv cs.CL TIER_1 English(EN) · Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah, Abdelrahman Eldesokey, Abeer Kashar, Abolade Daud, Abosede Grace Olanihun, Adamu Labaran Mohammed, Adeyemi Praise, Adhikarimayum Meerajita Sharma, Aditi Gupta, Adril Putra Merin, Adwoa Bremang, Afit… ·

    Global PIQA:跨越100多种语言和文化的常识推理评估

    arXiv:2510.24081v2 Announce Type: replace Abstract: To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense re…

  253. arXiv cs.CL TIER_1 English(EN) · Liang Chen, Xueting Han, Li Shen, Jing Bai, Kam-Fai Wong ·

    超越两阶段训练:LLM推理的协同SFT与RL

    arXiv:2509.06948v3 Announce Type: replace Abstract: Supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) are two widely used post-training paradigms for improving the reasoning ability of large language models (LLMs). Recent methods attempt to in…

  254. arXiv cs.CL TIER_1 English(EN) · Sharath Sathish ·

    Pramana:通过 Navya-Nyaya 对大型语言模型进行知识推理微调

    arXiv:2604.04937v1 Announce Type: cross Abstract: Large language models produce fluent text but struggle with systematic reasoning, often hallucinating confident but unfounded claims. When Apple researchers added irrelevant context to mathematical problems, LLM performance degrad…

  255. arXiv cs.CL TIER_1 English(EN) · Shashi Kumar, Yacouba Kaloga, Petr Motlicek, Ina Kodrasi, Andrea Cavallaro ·

    几何潜在推理促使大型语言模型生成更短的内容

    arXiv:2606.02248v1 Announce Type: new Abstract: Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive, length-sensitive, and constrained to (discrete) natural language. While latent …

  256. arXiv cs.CL TIER_1 English(EN) · Chengtao Gan, Zhiqiang Liu, Long Jin, Yushan Zhu, Lei Liang, Wen Zhang ·

    CRAFTQA:一种用于复杂结构化数据推理的代码驱动自适应框架

    arXiv:2606.02170v1 Announce Type: new Abstract: Real-world scenarios involve massive heterogeneous structured data (e.g., tables, knowledge graphs), making effective reasoning over such diverse data increasingly important. Unified structured data question answering has emerged as…

  257. arXiv cs.CL TIER_1 English(EN) · Ahmed Elhady, Eneko Agirre, Mikel Artetxe ·

    面向语言模型的跨语言一致性多语言推理

    arXiv:2606.01464v1 Announce Type: new Abstract: Despite expanding their multilingual coverage, the advanced reasoning capabilities of LLMs remain largely confined to a few high-resource languages like English. To address this, we propose an unsupervised Reinforcement Learning (RL…

  258. arXiv cs.CL TIER_1 English(EN) · Mengmeng Ji, Ravi Shanker Raju, Jonathan Lingjie Li, Chen Wu ·

    LongAttnComp: 跨家族上下文压缩以实现长上下文推理

    arXiv:2606.01336v1 Announce Type: new Abstract: As real-world applications increasingly require processing inputs of 100k+ tokens, the gap between context length and inference efficiency has become a critical bottleneck. Context compression offers a way to reduce prefill costs wh…

  259. arXiv cs.CL TIER_1 English(EN) · Ruiqi Zhang, Lingxiang Wang, Hainan Zhang Zhiming Zheng ·

    通过动态令牌选择实现鲁棒推理以进行分布对齐的自蒸馏

    arXiv:2606.00628v1 Announce Type: new Abstract: Self-distillation improves learning efficiency by rewriting reference answers as training data that better matches the model's own distribution. However, reference answers also introduce strong stylistic biases, causing the generati…

  260. arXiv cs.AI TIER_1 English(EN) · Arip Asadulaev, Rayan Banerjee, Fakhri Karray, Martin Takac ·

    TRM 中的潜在推理秘密上是一个策略改进算子

    arXiv:2511.16886v5 Announce Type: replace-cross Abstract: Recently, small models with latent recursion have obtained promising results on complex reasoning tasks. These results are typically explained by the theory that such recursion increases a networks depth, allowing it to co…

  261. arXiv cs.AI TIER_1 English(EN) · Yoonjeon Kim, Doohyuk Jang, Eunho Yang ·

    通过推理模型中的预测奖励来验证元认知

    arXiv:2510.03259v2 Announce Type: replace-cross Abstract: Recent research on reasoning models explores the meta-awareness of language models, including their ability to determine optimal thinking duration, recognize knowledge boundaries, and structure concept-level thinking. Whil…

  262. arXiv cs.AI TIER_1 English(EN) · Jiwoong Sohn, Tomasz Sternal, Kenneth Styppa, Torsten Hoefler, Michael Moor ·

    Process Reward Agents for Steering Knowledge-Intensive Reasoning

    arXiv:2604.09482v2 Announce Type: replace Abstract: Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating step correctness may require synthesizing clues across large external knowledge…

  263. arXiv cs.AI TIER_1 English(EN) · Nearchos Potamitis, Vansh Ramani, Har Ashish Arora, Dhairya Kuchhal, Lars Klein, Akhil Arora ·

    ReasonBENCH:衡量LLM推理的(不)稳定性

    arXiv:2512.07795v2 Announce Type: replace Abstract: Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different answers and costs across repeated executions, even under greedy decoding (T = 0…

  264. arXiv cs.AI TIER_1 English(EN) · Yaoming Li, Guangxiang Zhao, Qilong Shi, Lin Sun, Xiangzheng Zhang, Tong Yang ·

    训练后推理数据入门:我们对它的工作原理了解多少

    arXiv:2606.02113v1 Announce Type: cross Abstract: Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly,…

  265. arXiv cs.AI TIER_1 English(EN) · Dhruv Saini, Rohan Pandey ·

    ThinkSwitch:使用 LoRA 和权重插值进行特定目的推理任务的上下文蒸馏

    arXiv:2606.01080v1 Announce Type: cross Abstract: Large language models often improve on difficult tasks by spending inference-time compute on a reasoning trace before producing the final answer. That extra computation can be useful, but it also raises latency, token cost, and de…

  266. arXiv cs.AI TIER_1 English(EN) · Zihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding, Yining Sun ·

    结果优化悖论:LLM推理捷径的因果信息论界限

    arXiv:2606.00674v1 Announce Type: cross Abstract: Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reasoning capab…

  267. arXiv cs.AI TIER_1 English(EN) · Jiafu Huang, Chao Peng, Chenyang Xu, Zhengfeng Yang, Kecheng Cai, Chenhao Zhang, Yi Wang, Yiwei Gong, Wanqin Zhou, Irene Zheng ·

    通过辅助重建为神经算法推理提供更丰富的表示

    arXiv:2606.00559v1 Announce Type: cross Abstract: Neural algorithmic reasoning has emerged as a popular research direction. It aims to train neural networks to mimic the step-by-step behavior of classical rule-based algorithms. More specifically, the execution of such algorithms …

  268. arXiv cs.AI TIER_1 English(EN) · Ekaterina Alimaskina, Darya Rudas, Denis Shveykin, Gleb Molodtsov, Pavel Vasiliev, Aleksandr Beznosikov ·

    推理模型中的极低比特推理:失效模式与定向恢复

    arXiv:2606.02011v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) rely on long reasoning traces, making inference expensive. While low-bit quantization reduces per-token decoding cost, we show that aggressive 2-bit inference can fail to deliver end-to-end speedup beca…

  269. arXiv cs.AI TIER_1 English(EN) · Shayan Shokri ·

    TERRA:跨领域应用的任务嵌入式推理与表示架构

    arXiv:2606.01520v1 Announce Type: new Abstract: A single action-conditioned latent predictive architecture can in principle be trained on the structured state of a driving scene, a robot workspace, or a financial order book. The ingredients for doing so within any one domain alre…

  270. arXiv cs.AI TIER_1 English(EN) · Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan ·

    人工智能之谜:探究大型推理模型的生产-评估鸿沟

    arXiv:2606.01462v1 Announce Type: new Abstract: Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning t…

  271. arXiv cs.AI TIER_1 English(EN) · Teddy Ferdinan, Bart{\l}omiej Koptyra, Miko{\l}aj Langner, Tomasz Adamczyk, {\L}ukasz Radli\'nski, Maciej Markiewicz, Aleksander Szcz\k{e}sny, Stanis{\l}aw Wo\'zniak, Tymoteusz Romanowicz, Dzmitry Pihulski, Mateusz Zbrocki, Mateusz \'Smigielski, Micha{\l… ·

    Reasoning4Sciences:将推理语言模型连接到所有科学分支

    arXiv:2606.01145v1 Announce Type: new Abstract: While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow -- or lack of -- adoption of RLMs in other branches of …

  272. arXiv cs.AI TIER_1 English(EN) · Jiakang Li, Guanyu Zhu, Can Jin, Chenxi Huang, Dexu Yu, Ronghao Chen, Yang Zhou, Hongwu Peng, Xuanqi Lan, Dimitris N. Metaxas, Youhua Li ·

    Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

    arXiv:2606.00726v1 Announce Type: new Abstract: Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive…

  273. arXiv cs.AI TIER_1 English(EN) · Alessio Bruno ·

    AXIOM:一种以信任为先的神经符号执行架构,用于可验证的数学推理

    arXiv:2606.00671v1 Announce Type: new Abstract: We present AXIOM, a trust-first neuro-symbolic execution architecture for natural-language mathematical reasoning. In AXIOM, the language model functions strictly as a canonicalizer: it rewrites informal problem text into a narrow s…

  274. arXiv cs.AI TIER_1 English(EN) · Jayant Parashar, Suchendra M. Bhandarkar ·

    KACE:面向数学推理的知识自适应上下文工程

    arXiv:2606.00532v1 Announce Type: new Abstract: Context engineering can improve large language models without updating their weights, but mathematical reasoning exposes a key limitation: feedback accumulated in one growing prompt causes context bloat and limits the amount of lear…

  275. arXiv cs.AI TIER_1 English(EN) · Dongxin Guo, Jikun Wu, Siu Ming Yiu ·

    确定性地平线:当扩展推理失败,工具委托变得必要时

    arXiv:2606.00376v1 Announce Type: new Abstract: Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not due to preference biases, but limits rooted in the information-theoretic capacity of decoder-only attention. We establish: (1) an…

  276. arXiv cs.AI TIER_1 English(EN) · Shunchi Zhang, Jin Lu, Chuanyang Jin, Yichao Zhou, Zhining Zhang, Tianmin Shu ·

    MindZero:以零标注学习在线心智推理

    arXiv:2606.00240v1 Announce Type: new Abstract: Effective real-world assistance requires AI agents with robust Theory of Mind (ToM): inferring human mental states from their behavior. Despite recent advances, several key challenges remain, including (1) online inference with robu…

  277. arXiv cs.AI TIER_1 English(EN) · Mingyuan Fan, Weiguang Han, Daixin Wang, Cen Chen, Zhiqiang Zhang, Jun Zhou ·

    评估大型语言模型中的交互式推理:一个具有可执行游戏的层级基准

    arXiv:2606.00103v1 Announce Type: new Abstract: We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the task rules, must issue targeted queries to a hidden en…

  278. arXiv cs.AI TIER_1 English(EN) · Gregory Magarshak ·

    Grokers:基于自下而上的归纳理解和类型化知识图谱上的写入时智能

    arXiv:2606.00050v1 Announce Type: new Abstract: We present Grokers, an architecture for building persistent, structured comprehension of typed knowledge graphs through bottom-up inductive traversal of dependency subgraphs. Unlike retrieval-augmented generation (RAG), which pays f…

  279. Hugging Face Daily Papers TIER_1 English(EN) ·

    从智能体轨迹中诱导推理原语

    ReAct-style LLM agents often rediscover the same reasoning routines across problems, yet leave those routines trapped in transient scratchpads. We introduce Reasoning Primitive Induction, a single-pass method that mines successful ReAct traces, clusters recurrent reasoning moves,…

  280. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

    Agentic Chain-of-Thought Steering (ACTS) formulates reasoning steering as a Markov decision process to enable efficient, controllable chain-of-thought reasoning with token savings.

  281. Hugging Face Daily Papers TIER_1 English(EN) ·

    Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning

    Prompt-Level Distillation extracts reasoning patterns from teacher models to enhance student model performance while maintaining interpretability and reducing latency.

  282. arXiv cs.CL TIER_1 English(EN) · Andrea Cavallaro ·

    几何潜在推理促使大型语言模型生成更短的内容

    Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive, length-sensitive, and constrained to (discrete) natural language. While latent reasoning offers a continuous alternative, deter…

  283. arXiv cs.CL TIER_1 English(EN) · Wen Zhang ·

    CRAFTQA:用于复杂结构化数据推理的代码驱动自适应框架

    Real-world scenarios involve massive heterogeneous structured data (e.g., tables, knowledge graphs), making effective reasoning over such diverse data increasingly important. Unified structured data question answering has emerged as a prominent research trend, aiming to answer na…

  284. arXiv cs.AI TIER_1 English(EN) · Tong Yang ·

    训练后推理数据入门:我们对它的工作原理了解多少

    Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across data…

  285. Hugging Face Daily Papers TIER_1 English(EN) ·

    推理模型中的极低比特推断:失效模式与定向恢复

    Large Reasoning Models (LRMs) rely on long reasoning traces, making inference expensive. While low-bit quantization reduces per-token decoding cost, we show that aggressive 2-bit inference can fail to deliver end-to-end speedup because instability in the generation process inflat…

  286. arXiv cs.AI TIER_1 English(EN) · Yu Zhao, Hao Guan, Yongcheng Jing, Ying Zhang, Dacheng Tao ·

    MedCoG:通过元认知调控最大化医疗推理中的LLM推理密度

    arXiv:2602.07905v2 Announce Type: replace Abstract: Large Language Models (LLMs) have shown strong potential in complex medical reasoning yet face diminishing gains under inference scaling laws. While existing studies augment LLMs with various knowledge types, it remains unclear …

  287. arXiv cs.AI TIER_1 English(EN) · Arya Fayyazi, Mehdi Kamal, Massoud Pedram ·

    COFT:大型语言模型中公平性思维链推理的逆事实-一致性解码

    arXiv:2605.30641v1 Announce Type: cross Abstract: Large language models (LLMs) can reveal and amplify societal biases during chain-of-thought (CoT) generation. We present COFT (Chain of Fair Thought), a training-free decoding method that applies token-level fairness control at de…

  288. arXiv cs.AI TIER_1 English(EN) · Tom Pecher ·

    机器中的社会推理:探究大型语言模型辩论中的集体寻求真相动态

    arXiv:2605.30391v1 Announce Type: cross Abstract: Human reasoning has long been theorised to operate socially, not through isolated individual cognition, but through collective adversarial discourse, a framework known as the Argumentative Theory of Reasoning (ATR). Rather than re…

  289. arXiv cs.AI TIER_1 English(EN) · Saku Peltonen, August B{\o}gh R{\o}nberg, Andreas Plesner, Roger Wattenhofer ·

    GraphARC:用于基于图的抽象推理的综合基准测试

    arXiv:2605.31031v1 Announce Type: new Abstract: Relational reasoning lies at the heart of intelligence, but existing benchmarks are typically confined to formats such as grids or text. We introduce GraphARC, a benchmark for abstract reasoning on graph-structured data. GraphARC ge…

  290. arXiv cs.AI TIER_1 English(EN) · Tianrun Yu, Kaixiang Zhao, Chih-Chun Chen, Amanda Hughes, Taylor W. Killian, Fenglong Ma, Weitong Zhang, Porter Jenkins ·

    LARK:基于可学性轨迹选择的高效推理蒸馏

    arXiv:2605.30651v1 Announce Type: cross Abstract: We study trajectory selection for reasoning distillation, where teacher-generated reasoning trajectories are selectively used as supervision for a student model. Existing methods rely on heuristics such as trajectory quality or mo…

  291. arXiv cs.AI TIER_1 English(EN) · Archiki Prasad, Mandar Joshi, Kenton Lee, Mohit Bansal, Peter Shaw ·

    有效的推理链降低内在维度

    arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, yet the precise mechanisms by which different strategies facilitate generalizatio…

  292. arXiv cs.AI TIER_1 English(EN) · Elchanan Mossel ·

    可证伪性鸿沟:大型语言模型推理验证中的挑战

    arXiv:2601.02380v4 Announce Type: replace-cross Abstract: Recent reports claim that Large Language Models (LLMs) have achieved the ability to derive new science and exhibit human-level general intelligence. We argue that such claims are not rigorous scientific claims, as they do …

  293. arXiv cs.AI TIER_1 English(EN) · Yunhe Li, Hao Shi, Bowen Deng, Wei Wang, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Siyang Gao, Chao Wang, Shuang Qiu, Linqi Song ·

    学习利用洞察力进行非正式定理证明

    arXiv:2604.16278v2 Announce Type: replace Abstract: Although most of the automated theorem-proving approaches depend on formal proof systems, informal theorem proving can align better with large language models' (LLMs) strength in natural language processing. In this work, we ide…

  294. Hugging Face Daily Papers TIER_1 English(EN) ·

    几何潜在推理促使大型语言模型生成更短的内容

    Geometric Latent Reasoning formulates latent reasoning as a geometric path-approximation problem in pretrained token-embedding space, enabling continuous intermediate reasoning states that reduce generation length while maintaining accuracy.

  295. Hugging Face Daily Papers TIER_1 English(EN) ·

    LongAttnComp: 跨系列上下文压缩以实现长上下文推理

    LongAttnComp adapts AttnComp for long-context processing by fine-tuning lightweight attention layers and implementing token-level chunking and positional reordering techniques.

  296. Hugging Face Daily Papers TIER_1 English(EN) ·

    人工智能的谜团:探究大型推理模型中的生产-评估差距

    Large reasoning models exhibit a significant gap between their ability to produce and evaluate reasoning, with models showing answer confirmation bias that prevents accurate reasoning evaluation.

  297. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tianmin Shu ·

    MindZero:以零标注进行在线心智推理学习

    Effective real-world assistance requires AI agents with robust Theory of Mind (ToM): inferring human mental states from their behavior. Despite recent advances, several key challenges remain, including (1) online inference with robust uncertainty updates over multiple hypotheses;…

  298. arXiv cs.CL TIER_1 English(EN) · Yueyang Wang, Jiawei Fu, Baolong Bi, Xili Wang, Xiaoqing Liu ·

    HE-SNR:通过熵揭示潜在逻辑以指导 SWE-bench 的中期训练

    arXiv:2601.20255v3 Announce Type: replace-cross Abstract: SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabilities are fundamentally acquired during the mid-training phase and subsequently …

  299. arXiv cs.AI TIER_1 English(EN) · Xin Chen, Feng Jiang, Yiqian Zhang, Hardy Chen, Shuo Yan, Wenya Xie, Min Yang, Shujian Huang ·

    提问式推理:将推理大型语言模型从被动求解器转变为主动探究者

    arXiv:2601.22139v2 Announce Type: replace-cross Abstract: Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain fundamentally limited by a \emph{blind self-thinking} paradigm: performing extensive …

  300. arXiv cs.AI TIER_1 English(EN) · Jiayi Dai, Randy Goebel ·

    向理性主义者学习:提炼中间可解释的理由

    arXiv:2601.22531v2 Announce Type: replace-cross Abstract: Because of the pervasive use of deep neural networks (DNNs), especially in high-stakes domains, the interpretability of DNNs has received increased attention. The general idea of rationale extraction (RE) is to provide an …

  301. arXiv cs.AI TIER_1 English(EN) · Kiran Tomlinson, Tobias Schnabel, Adith Swaminathan, Jennifer Neville ·

    推理的推理:LLM中思维链代币复杂度的BAPO界限

    arXiv:2602.02909v2 Announce Type: replace Abstract: Inference-time scaling via chain-of-thought (CoT) reasoning is a major driver of state-of-the-art LLM performance, but it comes with substantial latency and compute costs. We address a fundamental theoretical question: how many …

  302. arXiv cs.AI TIER_1 English(EN) · Samuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb, Jiawei Li, Yibo Yang, Ebey Abraham, Sunando Sengupta, Eric Sommerlade, Michael Wooldridge, Philip Torr ·

    理解能力的边界基准测试

    arXiv:2602.14307v3 Announce Type: replace Abstract: As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improving, it will become increasingly hard for humans…

  303. arXiv cs.AI TIER_1 English(EN) · Yang Ouyang, Shuhang Lin, Jung-Eun Kim ·

    DenseSteer:引导小型语言模型进行密集数学推理

    arXiv:2605.29247v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong chain-of-thought (CoT) reasoning abilities, while smaller models (<= 3B parameters) significantly underperform on multi-step reasoning tasks. Based on empirical analyses of the Qwen-2.…

  304. arXiv cs.AI TIER_1 English(EN) · Yubo Li, Ramayya Krishnan, Rema Padman ·

    链条保持,答案折叠:对抗性压力下推理模型中的追踪-答案分离

    arXiv:2605.29087v1 Announce Type: new Abstract: Reasoning models are evaluated on single-turn benchmarks but deployed in multi-turn dialogue, where users push back on correct answers. Under sustained adversarial pressure we find a previously undocumented failure mode: the chain-o…

  305. arXiv cs.AI TIER_1 English(EN) · Pedro Orvalho, Marta Kwiatkowska, Guillem Aleny\`a, Felip Many\`a ·

    大型语言模型通过基于偏好的最大可满足性实现可靠推理

    arXiv:2605.29687v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at understanding natural language but struggle with optimisation tasks involving multiple constraints and user-defined preferences, which commonly arise in domains such as robotics. We propose a hy…

  306. arXiv cs.AI TIER_1 English(EN) · Venkat Akhil Lakkapragada ·

    CosmicFish-HRM:紧凑型语言模型中的分层循环机制自适应推理

    arXiv:2605.28919v1 Announce Type: cross Abstract: Large language models have achieved strong reasoning capabilities, though often at the cost of massive parameter counts and expensive inference. In this work, we explore a different direction: adaptive reasoning depth in compact l…

  307. arXiv cs.AI TIER_1 English(EN) · Nishal Thomas, Noel Thomas ·

    FormInv:数学推理基准中的语义不变性测量协议

    arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o from rank 2 to rank 4 and elevates Claude Haiku and DeepSeek V3 above it; these ran…

  308. arXiv cs.AI TIER_1 English(EN) · Shreyas Fadnavis, Praitayini Kanakaraj, Felix Wyss ·

    何时以及多久?读出-中介者视角在时间推理中的应用

    arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reasoning in language models, a $\sin$/$\cos$ probe recovers day-of-year from a layer…

  309. arXiv cs.AI TIER_1 English(EN) · Lukas Aichberger, Sepp Hochreiter ·

    解锁大型语言模型工作记忆以进行潜在推理

    arXiv:2605.30343v1 Announce Type: cross Abstract: To improve the reasoning capabilities of large language models, test-time compute is typically scaled by generating intermediate tokens before the final answer. However, this couples reasoning to autoregressive generation and ther…

  310. arXiv cs.AI TIER_1 English(EN) · G M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-Ghazaleh ·

    大型推理模型中的分层思维建模

    arXiv:2510.22437v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) solve complex tasks by generating long Chain-of-Thought (CoT) sequences; however, the emergent dynamics governing reasoning trajectories are not well understood and can lead to inconsistencies and r…

  311. arXiv cs.AI TIER_1 English(EN) · Xinyu Liu, Xin Liu, Bo Jin, Runsong Zhao, Pengcheng Huang, Junhao Ruan, Bei Li, Chunyang Xiao, Chenglong Wang, Tong Xiao, Jingbo Zhu ·

    MemoSight:统一上下文压缩与多令牌预测以加速推理

    arXiv:2604.14889v2 Announce Type: replace Abstract: While chain-of-thought (CoT) reasoning enables LLMs to solve challenging reasoning tasks, the linear growth of the KV cache leads to substantial memory and inference overhead. Existing approaches such as context compression and …

  312. arXiv cs.AI TIER_1 English(EN) · Zhicheng Yang, Zhijiang Guo, Yifan Song, Minrui Xu, Yongxin Wang, Yiwei Wang, Xiaodan Liang, Jing Tang ·

    Prune-OPD:高效可靠的同策略蒸馏用于长时推理

    arXiv:2605.07804v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models. However, scaling OPD to long-horizon tasks exposes a critical flaw: as the student's generated prefix inevitably diverges from the t…

  313. arXiv cs.CL TIER_1 English(EN) · Mayug Maniparambil, Arjun Karuvally, Terrence Sejnowski, Fergal Reid ·

    当强化学习抑制自身词汇:恢复拼图到数学迁移中的推理多样性

    arXiv:2605.29190v1 Announce Type: cross Abstract: Reinforcement learning using verifiable rewards (RLVR) improves LLM reasoning, but the conditions under which it transfers across domains -- and why it does so -- remain under-explored. We study cross-domain transfer in a 7B model…

  314. arXiv cs.CL TIER_1 English(EN) · Jun Rao, Zixiong Yu, Xuebo Liu, Guhan Chen, Jing Li, Hejin Wang, Jiansheng Wei, Xiaojun Meng, Min Zhang ·

    挖掘还是合成?重新思考数学推理迭代对齐中的探索效率

    arXiv:2602.05370v3 Announce Type: replace Abstract: Iterative Direct Preference Optimization (DPO) has emerged as a widely used paradigm for aligning Large Language Models on reasoning tasks. Existing approaches typically rely on Best-of-N sampling ($N\geq8$) to mine positive tra…

  315. arXiv cs.LG TIER_1 English(EN) · Jonathan Williams, Esin Tureci ·

    优先考虑过程,而非仅仅结果:奖励潜在思维轨迹可改善循环语言模型的推理能力

    arXiv:2602.10520v3 Announce Type: replace Abstract: Looped Language Models (LoopLMs) perform multi-step latent reasoning prior to token generation and outperform conventional LLMs on reasoning benchmarks at smaller parameter budgets. However, attempts to further improve LoopLM re…

  316. arXiv cs.CL TIER_1 English(EN) · Jia-Chen Zhang, Yu-Jie Xiong, Zheng Zhou ·

    认知循环思维:用于高效数学推理的可逆分层马尔可夫链

    arXiv:2604.06805v2 Announce Type: replace Abstract: Multi-step Chain-of-Thought (CoT) has significantly advanced the mathematical reasoning capabilities of LLMs by leveraging explicit reasoning steps. However, the widespread adoption of Long CoT often results in sequence lengths …

  317. Hugging Face Daily Papers TIER_1 English(EN) ·

    MindZero:以零标注进行在线心智推理学习

    MindZero presents a self-supervised reinforcement learning framework that enables multimodal large language models to perform efficient and robust online mental reasoning without requiring explicit mental state annotations.

  318. arXiv cs.AI TIER_1 English(EN) · Sepp Hochreiter ·

    解锁大型语言模型的工作记忆以进行潜在推理

    To improve the reasoning capabilities of large language models, test-time compute is typically scaled by generating intermediate tokens before the final answer. However, this couples reasoning to autoregressive generation and thereby conflates internal computation with external c…

  319. arXiv cs.AI TIER_1 English(EN) · Quanquan C. Liu ·

    基于采样进行推理:在决策点进行剪枝

    Frontier reasoning models are produced by posttraining base language models with reinforcement learning. Recent work has challenged this by showing that sampling from a sharpened version of the base model's distribution, a so-called power distribution, elicits comparable reasonin…

  320. arXiv cs.AI TIER_1 English(EN) · Guha Balakrishnan ·

    Conformal Certification of Reasoning Trace Prefixes

    Language model reasoning traces are rarely all-or-nothing; they frequently contain valid intermediate steps before a critical error occurs. Existing uncertainty quantification methods typically certify final answers or entire responses, failing to provide statistical guarantees f…

  321. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tom Pecher ·

    机器中的社会推理:探究大型语言模型辩论中的集体寻求真相动态

    Human reasoning has long been theorised to operate socially, not through isolated individual cognition, but through collective adversarial discourse, a framework known as the Argumentative Theory of Reasoning (ATR). Rather than relying on individual "intellectualist reasoners" as…

  322. arXiv cs.AI TIER_1 English(EN) · Xue Wen Tan, Nathaniel Tan, Galen Lee, Stanley Kok ·

    推理的形状:大型语言模型推理轨迹的拓扑分析

    arXiv:2510.20665v3 Announce Type: replace Abstract: Evaluating the quality of reasoning traces from large language models remains understudied, labor-intensive, and unreliable: current practice relies on expert rubrics, manual annotation, and slow pairwise judgments. Automated ef…

  323. arXiv cs.AI TIER_1 English(EN) · Linas Nasvytis, Simon Jerome Han, Ben Prystawski, Satchel Grant, Noah D. Goodman, Judith E. Fan ·

    CORE:对比性反思促进推理的快速改进

    arXiv:2605.28742v1 Announce Type: new Abstract: Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) approaches to doing so typically require hundreds of trai…

  324. arXiv cs.AI TIER_1 English(EN) · Biagio La Rosa, Leilani H. Gilpin ·

    神经元的保证最优组合解释

    arXiv:2511.20934v2 Announce Type: replace Abstract: Compositional explanations are a family of methods that aim to describe the spatial alignment between neurons' receptive field activations and concepts through logical rules, typically computed via a search over all possible con…

  325. arXiv cs.CL TIER_1 English(EN) · Shengmin Piao, Sanghyun Park ·

    GeneralThinker:通过似然引导的答案条件优化实现领域通用推理

    arXiv:2605.27934v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves language model reasoning, but its reliance on domain-specific verifiers, sparse outcome rewards, and coarse-grained credit assignment limits its applicability. We introduce Gen…

  326. arXiv cs.AI TIER_1 English(EN) · Guoxin Ma, Yibing Liu, Chengzhengxu Li, Yu Liang, Yan Wang, Yueyang Zhang, Kecheng Chen, Zhaohan Zhang, Zhiyuan Sun, Daiting Shi ·

    思考即压缩:你的推理模型秘密是一个上下文压缩器

    arXiv:2605.28713v1 Announce Type: new Abstract: Context compression aims to shorten long context inputs with minimal information loss for LLM inference acceleration. While existing methods have shown promise, they typically rely on complex compression modules or compression-speci…

  327. arXiv cs.AI TIER_1 English(EN) · Leizhen Zhang, Shuhan Chen, Sheng Chen ·

    LLM 满足可满足性求解:推理能力的匹配对评估

    arXiv:2605.28602v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks that implicitly reduce to Boolean satisfiability (SAT), yet their reasoning ability on SAT remains unclear. We present a systematic study of LLMs on 2-SAT and 3-SAT, toget…

  328. arXiv cs.CL TIER_1 English(EN) · Ziqi Zhao, Xinyu Ma, Liu Yang, Yujie Feng, Daiting Shi, Jingzhou He, Xin Xin, Zhaochun Ren, Xiao-Ming Wu ·

    ROSD: 跨领域语言模型推理的自反策略自蒸馏

    arXiv:2605.28014v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, existing OPSD methods often yield limited gains on in-…

  329. arXiv cs.CL TIER_1 English(EN) · Yukyung Lee, Yumeng Shen, Jinhyeong Park, Hyein Yang, Jun-Hyung Park ·

    CIRF:将思维链进行令牌化,转化为可重用的功能单元,以实现大型语言模型中高效的潜在推理

    arXiv:2605.28292v1 Announce Type: new Abstract: Implicit Chain-of-Thought (CoT) reduces the inference cost of large language models by internalizing the explicit rationales. However, existing approaches typically lack alignment with explicit rationales and adaptivity to example c…

  330. arXiv cs.LG TIER_1 English(EN) · Avidan Shah, Jannik Brinkmann, Rico Angell ·

    通过激活一致性训练减轻针对推理模型的自适应攻击

    arXiv:2605.28467v1 Announce Type: new Abstract: As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection. We study consistency training, a family of fine-tu…

  331. arXiv cs.AI TIER_1 English(EN) · Phuong Minh Nguyen, Tien Huu Dang, Naoya Inoue ·

    揭示用于逻辑推理的算法演绎电路

    arXiv:2605.27824v1 Announce Type: new Abstract: Recent studies have shown that Large Language Models (LLMs) can achieve strong reasoning performance by incorporating functional symbolic representations that abstractly describe graph traversal algorithms and step-by-step reasoning…

  332. arXiv cs.AI TIER_1 English(EN) · Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou Ammar ·

    风险可控的 Lean-as-Judge 用于自然语言数学推理

    arXiv:2605.28365v1 Announce Type: new Abstract: Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. …

  333. arXiv cs.AI TIER_1 English(EN) · Renjie Gu, Jiaxu Li, Yihao Wang, Yun Yue, Hansong Xiao, Yefei Chen, Yuan Wang, Chunxiao Guo, Pei Wei, Jinjie Gu, Yixin Cao ·

    在信息不足的情况下弥合推理模型中的检测与弃权差距

    arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified, yet still continue reasoning and produce unsupported final answers instead of …

  334. arXiv cs.AI TIER_1 English(EN) · Navid Rezazadeh, Arash Gholami Davoodi ·

    过度思考的形状:长推理轨迹中的回溯爆发

    arXiv:2605.27965v1 Announce Type: new Abstract: Reasoning models often generate long traces in which useful self-correction and unproductive revision are hard to distinguish. We study this distinction through backtracking dynamics: local reconsideration, retraction, or re-derivat…

  335. arXiv cs.AI TIER_1 English(EN) · Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo ·

    压缩思维:LLM 后训练中压缩推理数据何时以及如何奏效

    arXiv:2605.28008v1 Announce Type: new Abstract: Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and token cost remains a central challenge. To address this issue, supervised fine-tuni…

  336. arXiv cs.AI TIER_1 English(EN) · Chien-Ping Lu ·

    推理的计算边界:能力内化、训练与图灵跃迁

    arXiv:2605.27381v1 Announce Type: cross Abstract: Claims about recursive self-improvement in AI often slide from repeated internal revision to the possibility of qualitatively stronger capability without clearly distinguishing the underlying computational regimes. This paper give…

  337. arXiv cs.AI TIER_1 English(EN) · Taylor Olson, Roberto Salas-Damian, Kenneth D. Forbus ·

    动态变化的规范下的推理与规划

    arXiv:2605.27622v1 Announce Type: new Abstract: To safely interact with humans, AI agents must both know our norms and consider them during planning. However, such norm-guided planning has been less explored, only within communities of artificial agents, and has ignored the dynam…

  338. arXiv cs.AI TIER_1 English(EN) · Judith E. Fan ·

    CORE:对比性反思促进推理能力的快速提升

    Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) approaches to doing so typically require hundreds of training samples and thousands of model rollouts, ma…

  339. arXiv cs.AI TIER_1 English(EN) · Daiting Shi ·

    思考即压缩:你的推理模型其实是上下文压缩器

    Context compression aims to shorten long context inputs with minimal information loss for LLM inference acceleration. While existing methods have shown promise, they typically rely on complex compression modules or compression-specific training, leaving the intrinsic capabilities…

  340. arXiv cs.AI TIER_1 English(EN) · Sheng Chen ·

    LLM 满足可满足性求解:推理能力的匹配对评估

    Large language models (LLMs) are increasingly used for tasks that implicitly reduce to Boolean satisfiability (SAT), yet their reasoning ability on SAT remains unclear. We present a systematic study of LLMs on 2-SAT and 3-SAT, together with two canonical reductions, Vertex Cover …

  341. arXiv cs.LG TIER_1 English(EN) · Rico Angell ·

    通过激活一致性训练减轻针对推理模型的自适应攻击

    As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection. We study consistency training, a family of fine-tuning objectives that enforce identical behavior …

  342. arXiv cs.CL TIER_1 English(EN) · Haitham Bou Ammar ·

    风险可控的 Lean-as-Judge 用于自然语言数学推理

    Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. On MATH-500 we show this signal is (i) sharply c…

  343. arXiv cs.CL TIER_1 English(EN) · Jun-Hyung Park ·

    CIRF:将思维链分解为可重用的功能单元,以实现大型语言模型中高效的潜在推理

    Implicit Chain-of-Thought (CoT) reduces the inference cost of large language models by internalizing the explicit rationales. However, existing approaches typically lack alignment with explicit rationales and adaptivity to example complexity. In this work, we propose CIRF (\texti…

  344. Hugging Face Daily Papers TIER_1 English(EN) ·

    在信息不足的情况下弥合推理模型中的检测与弃权差距

    We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified, yet still continue reasoning and produce unsupported final answers instead of abstaining. We formalize this mismatch as the de…

  345. Hugging Face Daily Papers TIER_1 English(EN) ·

    ROSD: 跨领域语言模型推理的反射性策略内蒸馏

    On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, existing OPSD methods often yield limited gains on in-domain reasoning and generalize poorly to out-of…

  346. Hugging Face Daily Papers TIER_1 English(EN) ·

    压缩思维:LLM 后训练中压缩推理数据何时以及如何起作用

    Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and token cost remains a central challenge. To address this issue, supervised fine-tuning (SFT) often uses compressed reasoning data, w…

  347. Hugging Face Daily Papers TIER_1 English(EN) ·

    GeneralThinker:通过似然引导的答案条件优化实现领域通用推理

    Reinforcement learning with verifiable rewards improves language model reasoning, but its reliance on domain-specific verifiers, sparse outcome rewards, and coarse-grained credit assignment limits its applicability. We introduce GeneralThinker, an on-policy framework that reformu…

  348. arXiv cs.CL TIER_1 English(EN) · Lisong Sun, Li Wang, Chen Zhang, Jinyang Wu, Kui Zhang, Tianhao Peng, Wenjun Wu ·

    学习自适应SFT数据以提升推理泛化能力

    arXiv:2605.26924v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Among post-training paradigms, supervised fine-tuning (SFT) is widely used: it leve…

  349. arXiv cs.AI TIER_1 English(EN) · Seonghoon Yu, Dongjun Nam, Byung-Kwan Lee, Jeany Son ·

    隐藏以见:视觉锚定思维中的VLM蒸馏推理前缀掩码

    arXiv:2605.11651v4 Announce Type: replace-cross Abstract: Recent think-answer approaches in VLMs, such as Qwen3-VL-Thinking, boost reasoning performance by leveraging intermediate thinking steps before the final answer, but their computational cost becomes substantial, especially…

  350. arXiv cs.AI TIER_1 English(EN) · Xuhang Chen, Zhifan Song, Deyi Ji, Shuo Gao, Lanyun Zhu ·

    自信号驱动的多LLM辩论,实现高效准确推理

    arXiv:2510.06843v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have exhibited impressive capabilities across diverse application domains. Recent work has explored Multi-LLM Agent Debate (MAD) as a way to enhance performance by enabling multiple LLMs to dis…

  351. arXiv cs.AI TIER_1 English(EN) · Hans Peter Lyngs{\o}e Raaschou-Jensen, Constanza Fierro, Anders S{\o}gaard ·

    Reasoning Language Models 的实时进度预测

    arXiv:2506.23274v4 Announce Type: replace-cross Abstract: Recent reasoning language models, particularly those that employ long latent chains of thought, achieve strong performance on complex agentic tasks. However, as these models operate over increasingly long time horizons, th…

  352. arXiv cs.AI TIER_1 English(EN) · Meghyn Bienvenu, Camille Bourgaux ·

    查询和修复不一致的优先知识库:复杂性分析与抽象论证的联系

    arXiv:2003.05746v4 Announce Type: replace-cross Abstract: In this paper, we explore the issue of inconsistency handling over prioritized knowledge bases (KBs), which consist of an ontology, a set of facts, and a priority relation between conflicting facts. In the database setting…

  353. arXiv cs.AI TIER_1 English(EN) · Yihua Zhu, Qianying Liu, Fei Cheng, Jiaxin Wang, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira ·

    推理深度与环境复杂度:RLVR数据分配在逻辑推理任务上的受控研究

    arXiv:2605.26934v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become central to post-training reasoning models, yet a key limitation of existing studies is their narrow view of the reasoning space: difficulty is treated as reasoning d…

  354. arXiv cs.LG TIER_1 English(EN) · Alex Ayoub, Kavosh Asadi, Dale Schuurmans, Csaba Szepesv\'ari, Karim Bouyarmane ·

    使用折扣强化学习高效学习推理

    arXiv:2510.23486v2 Announce Type: replace Abstract: Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to reach the goal quickly, and LRM reasoning can b…

  355. arXiv cs.AI TIER_1 English(EN) · Xiao-Wen Yang, Ziyu Han, Xi-Hua Zhang, Wen-Da Wei, Jie-Jing Shao, Lan-Zhe Guo, Yu-Feng Li ·

    为循环语言模型中的测试时可扩展潜在推理稳定循环动力学

    arXiv:2605.26733v1 Announce Type: cross Abstract: Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: performance often peaks at a certain iteration depth and then collapses with further r…

  356. arXiv cs.AI TIER_1 English(EN) · Shanghao Li, Jinda Han, Yibo Wang, Yuanjie Zhu, Zihe Song, Langzhou He, Kenan Kamel A Alghythee, Philip S. Yu ·

    大型语言模型为何会在结构化知识上产生幻觉:线性化表示的推理机制分析

    arXiv:2605.26362v1 Announce Type: cross Abstract: In many reasoning tasks, large language models (LLMs) rely on structured external knowledge, such as graphs and tables, which is typically linearized into sequential token representations. However, even when sufficient knowledge i…

  357. arXiv cs.AI TIER_1 English(EN) · Zhe Yu, Wenpeng Xing, Yunzhao Wei, Jie Chen, Hongzhi Wang, Xuyang Teng, Meng Han ·

    组合崩溃:稳定的事实知识并不意味着组合推理能力

    arXiv:2605.26789v1 Announce Type: new Abstract: Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly must be better at assembling facts. We show that th…

  358. Hugging Face Daily Papers TIER_1 English(EN) ·

    揭示用于逻辑推理的算法演绎电路

    Recent studies have shown that Large Language Models (LLMs) can achieve strong reasoning performance by incorporating functional symbolic representations that abstractly describe graph traversal algorithms and step-by-step reasoning in few-shot learning settings. However, it rema…

  359. Hugging Face Daily Papers TIER_1 English(EN) ·

    链条依旧,答案折叠:对抗性压力下推理模型的追踪-答案分离

    Research reveals a new failure mode in reasoning models where correct chain-of-thought reasoning leads to incorrect final answers under adversarial conditions, demonstrated through controlled experiments across multiple datasets and models.

  360. Hugging Face Daily Papers TIER_1 English(EN) ·

    揭示用于逻辑推理的算法演绎电路

    Large language models use specialized attention heads for retrieving factual information and integrating multi-step reasoning, with distinct neural mechanisms for local reasoning steps versus global strategy coordination.

  361. Hugging Face Daily Papers TIER_1 English(EN) ·

    CORE:对比反思促进推理能力的快速提升

    Contrastive Reflection (CORE) improves language model reasoning by analyzing differences between successful and unsuccessful attempts to generate concise, interpretable insights that enable faster and more efficient self-improvement compared to traditional parametric and non-para…

  362. arXiv cs.AI TIER_1 English(EN) · Hidetoshi Shimodaira ·

    推理深度与环境复杂度:RLVR数据分配在逻辑推理任务上的对照研究

    Reinforcement learning with verifiable rewards (RLVR) has become central to post-training reasoning models, yet a key limitation of existing studies is their narrow view of the reasoning space: difficulty is treated as reasoning depth alone, and reward is concentrated on forward …

  363. arXiv cs.CL TIER_1 English(EN) · Wenjun Wu ·

    学习适应SFT数据以获得更好的推理泛化能力

    Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Among post-training paradigms, supervised fine-tuning (SFT) is widely used: it leverages external data to provide dense supervision…

  364. arXiv cs.CL TIER_1 English(EN) · Yuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang ·

    哪些推理轨迹能教会学生更好地推理?信息对齐的简单指标

    arXiv:2601.14249v5 Announce Type: replace Abstract: Long chain-of-thought (CoT) trajectories provide rich supervision signals for distilling reasoning from teacher to student LLMs. However, both prior work and our experiments show that trajectories from stronger teachers do not n…

  365. arXiv cs.LG TIER_1 English(EN) · Wenbo Pan, Zhichao Liu, Xianlong Wang, Haining Yu, Xiaohua Jia ·

    迈向长时域可解释性:面向推理大语言模型的、高效且忠实的、多Token归因方法

    arXiv:2602.01914v2 Announce Type: replace Abstract: Token attribution methods provide intuitive explanations for language model outputs by identifying causally important input tokens. However, as modern LLMs increasingly rely on extended reasoning chains, existing schemes face tw…

  366. arXiv cs.CL TIER_1 English(EN) · Lisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey, Hossein A. Rahmani, Marek Rei ·

    AgentCoMa:一个在真实场景中混合了常识和数学推理的组合式基准测试

    arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple reasoning steps. However, current compositional benchmarks testing these skills t…

  367. arXiv cs.CL TIER_1 English(EN) · Hui Xie, Jie Liu, Ziyue Qiao, Joaquin Vanschore ·

    选择性潜在思维:LLM推理链的自适应压缩

    arXiv:2605.25745v1 Announce Type: new Abstract: Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost due to lengthy autoregressive traces. Existing latent reasoning methods offer a p…

  368. arXiv cs.CL TIER_1 English(EN) · Zongji Yu, Wenshui Luo, Yiliu Sun, Hao Fang, Runmin Cong, Chaochao Lu, Chen Gong ·

    多元协同:面向大型推理模型的跨域对比策略优化

    arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimization (GRPO). However, GRPO-style RL methods in multi-…

  369. arXiv cs.CL TIER_1 Norsk(NO) · Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, Leszek Rutkowski, Dacheng Tao ·

    更优、更快:利用大型推理模型的自我改进能力

    arXiv:2605.24998v1 Announce Type: new Abstract: Self-improvement training enables the large reasoning models (LRMs) to improve themselves by self-generating reasoning trajectories as training data without external supervision. However, we find that this method often falls short i…

  370. arXiv cs.AI TIER_1 English(EN) · Serafim Batzoglou ·

    归纳:一阶逻辑中的有限结构概念综合

    arXiv:2602.18956v3 Announce Type: replace Abstract: We introduce INDUCTION, a benchmark for finite structure concept synthesis in first order logic. Given small finite relational worlds with extensionally labeled target predicates, models must output a single first order logical …

  371. arXiv cs.AI TIER_1 English(EN) · Szymon Bobek, {\L}ukasz Ba{\l}ec, Grzegorz J. Nalepa ·

    结合领域知识和合理性约束的可操作且多样化的反事实解释

    arXiv:2511.20236v3 Announce Type: replace Abstract: Counterfactual explanations improve the actionable interpretability of machine learning models by identifying minimal changes required to achieve a desired outcome. However, existing methods often neglect dependencies among feat…

  372. arXiv cs.AI TIER_1 English(EN) · Mingyu Zhang, Lifeng Zhuo, Tianxi Tan, Guocan Xie, Xian Nie, Yan Li, Renjie Zhao, Zizhu He, Ziyu Wang, Jiting Cai, Yong-Lu Li ·

    IPR-1:交互式物理推理器

    arXiv:2511.15407v4 Announce Type: replace Abstract: Humans learn by observing, interacting with environments, and internalizing physics and causality. Here, we aim to ask whether an agent can similarly acquire human-like reasoning from interaction and keep improving with more exp…

  373. arXiv cs.AI TIER_1 English(EN) · Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur, Jason Gusdorf, Sourik Beltr\'an, Bita Behrouzi, Byron Crowe, Jacob Dockterman, Muzzammil Muhammad, Sarah Ohnigian, Andrew Sanchez, James A. Diao, Aashna P. Shah, Daniel Restrepo, Eric S. Rosenberg, And… ·

    教会大型语言模型像专家诊断师一样推理

    arXiv:2509.12194v2 Announce Type: replace Abstract: Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, fe…

  374. arXiv cs.AI TIER_1 English(EN) · Qirun Dai, Xiao Liu, Jiawei Zhang, Dylan Zhang, Hao Peng, Chenhao Tan ·

    迈向通用因果推理器

    arXiv:2605.24873v1 Announce Type: cross Abstract: Despite the importance of causal reasoning, training LLMs to reason causally remains underexplored. Existing data efforts mostly focus on benchmarking LLMs on specific aspects of causality, making them less suitable for training g…

  375. arXiv cs.AI TIER_1 English(EN) · Hongbo Jin, Mingnan Zhu, Jingqi Tian, Xu Jiang, Zhongjing Du, Haoran Tang, Siyi Xie, Qiaoman Zhang, Jiayu Ding ·

    Context-CoT:通过高质量推理合成增强上下文学习

    arXiv:2605.25354v1 Announce Type: new Abstract: While LLMs excel at reasoning over prompts using static pretrained knowledge, they struggle significantly with context learning-the ability to dynamically extract, internalize, and apply new knowledge from complex, task-specific con…

  376. arXiv cs.AI TIER_1 English(EN) · Andrew Corbett, Archit Sood, Anna Tzatzopoulou, Sai-Aakash Ramesh, Tim Dodwell ·

    通过引导式推理提升推理性能:递归模型的随机探索

    arXiv:2605.25230v1 Announce Type: new Abstract: Recent work on recursive architectures has shown that tiny neural networks can be surprisingly powerful on structured reasoning tasks. The trick is to model reasoning trajectories with a latent dynamical system. We argue that the in…

  377. arXiv cs.AI TIER_1 English(EN) · Andreas Opedal, Francesco Ignazio Re, Abulhair Saparov, Mrinmaya Sachan, Bernhard Sch\"olkopf, Ryan Cotterell ·

    通过 A* 训练后学习高效推理

    arXiv:2605.24597v1 Announce Type: new Abstract: Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We frame natural language inference as a search problem where the final answer is t…

  378. Hugging Face Daily Papers TIER_1 English(EN) ·

    选择性潜在思维:LLM推理链的自适应压缩

    Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost due to lengthy autoregressive traces. Existing latent reasoning methods offer a promising alternative, yet they often treat reaso…

  379. arXiv cs.CL TIER_1 English(EN) · Joaquin Vanschore ·

    选择性潜在思维:LLM推理链的自适应压缩

    Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost due to lengthy autoregressive traces. Existing latent reasoning methods offer a promising alternative, yet they often treat reaso…

  380. arXiv cs.CL TIER_1 English(EN) · Chen Gong ·

    多元协同:面向大型推理模型的跨域对比策略优化

    Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimization (GRPO). However, GRPO-style RL methods in multi-domain settings often fail to achieve consistent…

  381. arXiv cs.LG TIER_1 English(EN) · Meir Roketlishvili, Semyon Semenov, Maksim Bobrin, Viktor Kovalchuk, Albert Baichorov, Abduragim Shtanchaev, Fakhri Karray, Dmitry V. Dylov, Martin Tak\'a\v{c}, Arip Asadulaev ·

    Convex Compositional Reasoning Models

    arXiv:2605.23395v1 Announce Type: new Abstract: Compositional energy-based models can generalize to larger combinatorial reasoning problems by reusing a learned factor energy across many local constraints. In our paper, we show that a key bottleneck in compositional reasoning is …

  382. arXiv cs.LG TIER_1 English(EN) · Hoang Phan, Quang H. Nguyen, Hung T. Q. Le, Xiusi Chen, Heng Ji, Khoa D. Doan ·

    大型推理模型中的批判机制解码

    arXiv:2603.16331v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) exhibit backtracking and self-verification mechanisms that enable them to revise intermediate steps and reach correct solutions, yielding strong performance on complex logical benchmarks. We hypothe…

  383. arXiv cs.CL TIER_1 English(EN) · Zhe Yuan, Yipeng Zhou, Jinghan Li, Xinyuan Chen, Bowen Deng, Zhiqian Chen, Liang Zhao ·

    LambdaPO:一种用于推理语言模型的Lambda风格策略优化

    arXiv:2605.19416v2 Announce Type: replace Abstract: Group Relative Policy Optimization(GRPO) has become a cornerstone of modern reinforcement learning alignment, prized for its efficacy in foregoing an explicit value-critic by leveraging reward normalization across sampled trajec…

  384. arXiv cs.AI TIER_1 English(EN) · Junyao Yang, Chen Qian, Kun Wang, Linfeng Zhang, Quanshi Zhang, Yong Liu, Dongrui Liu ·

    熵梯度反演:迈向大型推理模型的内部机制

    arXiv:2605.17770v2 Announce Type: replace Abstract: The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive ``fast thinking'' text generation to systematic, step-by-step ``slow thinking'' reasoning, unlocking state-of-the-art performance in c…

  385. Hugging Face Daily Papers TIER_1 English(EN) ·

    看得更多就意味着知道得更多吗?多源视觉推理的单锚定优势归一化

    A novel mono-anchored multi-source reasoning framework that uses dynamic anchors to quantify information gain and regulate modality interactions during reinforcement learning with verifiable rewards.

  386. arXiv cs.LG TIER_1 English(EN) · Arip Asadulaev ·

    Convex Compositional Reasoning Models

    Compositional energy-based models can generalize to larger combinatorial reasoning problems by reusing a learned factor energy across many local constraints. In our paper, we show that a key bottleneck in compositional reasoning is not composition itself, but the non-convex geome…

  387. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型推理模型中的批判机制解码

    Large Reasoning Models demonstrate hidden critique abilities that allow error recovery through internal mechanisms, identified via interpretable critique vectors that enhance error detection without additional training.

  388. Hugging Face Daily Papers TIER_1 English(EN) ·

    Equilibrium Reasoners: 学习吸引子实现可扩展推理

    Equilibrium Reasoners enable scalable reasoning through task-conditioned attractors that guide latent dynamical systems toward valid solutions, achieving significant accuracy improvements through iterative test-time computation.

  389. Hugging Face Daily Papers TIER_1 English(EN) ·

    LambdaPO:一种用于推理语言模型的Lambda风格策略优化

    Group Relative Policy Optimization(GRPO) has become a cornerstone of modern reinforcement learning alignment, prized for its efficacy in foregoing an explicit value-critic by leveraging reward normalization across sampled trajectory cohorts. However, the method's reliance on a mo…

  390. arXiv cs.CV TIER_1 English(EN) · Hong Yang, Basura Fernando ·

    ERQA-Plus:具身智能推理的诊断基准

    arXiv:2606.17639v1 Announce Type: cross Abstract: Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observa…

  391. arXiv cs.CV TIER_1 English(EN) · Basura Fernando ·

    ERQA-Plus:具身智能推理的诊断基准

    Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observations. Yet existing visual and embodied question a…

  392. arXiv cs.CV TIER_1 English(EN) · Chaoyu Li, Deeparghya Dutta Barua, Fei Tao, Pooyan Fazli ·

    CASHEW:通过迭代轨迹聚合稳定多模态推理

    arXiv:2601.08010v2 Announce Type: replace Abstract: Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sampling over the same input often produces diverge…

  393. arXiv stat.ML TIER_1 English(EN) · Ousmane Amadou Dia ·

    面向长篇推理的自适应核截断

    arXiv:2606.13982v1 Announce Type: new Abstract: Sampling plays an important role in long-form language-model reasoning. Over thousands of decoding steps, small changes in the candidate token set can compound into different reasoning trajectories, stability profiles, and final ans…

  394. arXiv stat.ML TIER_1 English(EN) · Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong ·

    断裂的思维链推理

    arXiv:2505.12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining. Similarly, Chain-of-T…

  395. arXiv stat.ML TIER_1 English(EN) · Ousmane Amadou Dia ·

    面向长篇推理的自适应核截断

    Sampling plays an important role in long-form language-model reasoning. Over thousands of decoding steps, small changes in the candidate token set can compound into different reasoning trajectories, stability profiles, and final answers. Existing truncation methods such as top-$p…

  396. arXiv cs.CV TIER_1 English(EN) · Han Huang, Hao Wang, Mengqi Zhang, Shu Wu, Qiang Liu, Liang Wang ·

    CRANE:用于推理多模态大语言模型的知识编辑

    arXiv:2606.09033v1 Announce Type: new Abstract: The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought (CoT) reasoning before producing answers, has introduced a new challenge for knowledge editing: methods that appear succes…

  397. arXiv cs.CV TIER_1 English(EN) · Leyi Wu, Yifan Zhao, Jinjie Zhang, Yinchuan Li, Ying-Cong Chen ·

    正确的推理策略就够了:EgoCross挑战赛的近乎无训练的领域推理

    arXiv:2606.00829v1 Announce Type: new Abstract: EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from surgery, industrial assembly, extreme sports, and animal-mounted cameras rather t…

  398. arXiv stat.ML TIER_1 English(EN) · Felix Zhou, Anay Mehrotra, Quanquan C. Liu ·

    采样推理:在决策点进行剪枝

    arXiv:2605.30327v1 Announce Type: cross Abstract: Frontier reasoning models are produced by posttraining base language models with reinforcement learning. Recent work has challenged this by showing that sampling from a sharpened version of the base model's distribution, a so-call…

  399. arXiv stat.ML TIER_1 English(EN) · Matt Y. Cheung, Ashok Veeraraghavan, Hanjie Chen, Guha Balakrishnan ·

    Conformal Certification of Reasoning Trace Prefixes

    arXiv:2605.30085v1 Announce Type: cross Abstract: Language model reasoning traces are rarely all-or-nothing; they frequently contain valid intermediate steps before a critical error occurs. Existing uncertainty quantification methods typically certify final answers or entire resp…

  400. arXiv cs.CV TIER_1 English(EN) · Fanhu Zeng, Zhicong Luo, Zefan Wang, You Li, Chi Chen, Maosong Sun ·

    看得更多就意味着知道得更多吗?多源视觉推理的单锚点优势归一化

    arXiv:2605.25437v1 Announce Type: new Abstract: Visual reasoning through reinforcement learning with verifiable rewards (RLVR) has achieved remarkable progress. However, when dealing with multi-source inputs, existing approaches tend to treat them as a mere accumulation of inform…

  401. Together AI blog TIER_1 English(EN) ·

    大型推理模型在推理过程中无法遵循指令:一项基准研究

    ReasonIF finds frontier LRMs fail to follow reasoning instructions >75% of the time; introduces a benchmark across languages, formatting, and length.

  402. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    VibeThinker-3B:基于Qwen2.5-Coder-3B和Spectrum-to-Signal后训练流水线的3B密集推理模型

    <p>VibeThinker-3B, a 3B MIT-licensed reasoning model matching DeepSeek V3.2 and Kimi K2.5 on verifiable benchmarks.</p> <p>The post <a href="https://www.marktechpost.com/2026/06/19/vibethinker-3b-a-3b-dense-reasoning-model-built-on-qwen2-5-coder-3b-with-the-spectrum-to-signal-pos…

  403. Medium — fine-tuning tag TIER_1 English(EN) · Dave R - Microsoft Azure & AI MVP☁️ ·

    微软Foundry中开源推理模型的训练后优化:从生产痕迹到…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/codex/post-training-open-source-reasoning-models-in-microsoft-foundry-from-production-traces-to-a-0362349438a0?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/…

  404. Towards AI TIER_1 English(EN) · Nehdiii ·

    强化学习能否帮助LLM发现新的推理策略?

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/can-reinforcement-learning-help-llms-discover-new-reasoning-strategies-f50b1b054ec7?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1790/0*dn25jRvK-xOFGOd6.…

  405. dev.to — MCP tag TIER_1 English(EN) · curatedmcp ·

    Sequential Thinking MCP:将难题分解为可解决的步骤

    <blockquote> <p><em>Install guide and config at <a href="https://curatedmcp.com/install/sequential-thinking-mcp/claude-desktop" rel="noopener noreferrer">curatedmcp.com</a></em></p> </blockquote> <h1> Sequential Thinking MCP: Break Down Hard Problems Into Solvable Steps </h1> <p>…

  406. Towards AI TIER_1 English(EN) · Faheem Munshi ·

    思维链提示: 让 AI 按部就班地推理 — 从提示到盈利 · 30天中的第11天

    <h4><em>The single technique that separates AI users who get plausible answers from those who get genuinely intelligent ones.</em></h4><p>Welcome to Week 3. For the past two weeks, you’ve been building your foundation — prompting structure, templates, roles, workflows. Today we s…

  407. Medium — Claude tag TIER_1 English(EN) · Chris Jones ·

    AI Weldr 助力确定性逻辑表达民主化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jzace42011/democratizing-the-expression-of-deterministic-logic-with-ai-weldr-2d2699db4578?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/862/1*VS1Lu1HNhBVyG5RGbBS0Pw.p…

  408. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    o1和DeepSeek-R1等推理模型通过在推理时生成显式的思维链与标准LLM不同——其架构是这样的

    Reasoning models like o1 and DeepSeek-R1 differ from standard LLMs by generating an explicit chain of thought at inference time — here is how that architecture actually works. https://www. nerdheadz.com/blog/reasoning-m odels-explained-o1-deepseek-r1-rlms # ai # machinelearning

  409. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    VibeThinker-3B:基于Qwen2.5-Coder-3B和Spectrum-to-Signal后训练流水线构建的3B密集推理模型VibeThinker-3B,一个3B MIT许可的推理

    VibeThinker-3B: A 3B Dense Reasoning Model Built on Qwen2.5-Coder-3B With the Spectrum-to-Signal Post-Training Pipeline VibeThinker-3B, a 3B MIT-licensed reasoning model matching DeepSeek V3.2 and ... #AI #Paper #Summary #AI #Shorts #Applications #Artificial #Intelligence #Editor…

  410. dev.to — LLM tag TIER_1 English(EN) · Michael "Mike" K. Saleme ·

    当安全护栏变成攻击目标:针对大语言模型安全层的推理扩展拒绝服务攻击

    <p>New research from HKUST (<a href="https://arxiv.org/abs/2606.14517" rel="noopener noreferrer">arXiv:2606.14517</a>, June 12) turns the agent safety layer into the attack surface.</p> <h2> What happened </h2> <p>Reasoning-based guardrails — the LLM safety layers that screen an …

  411. r/MachineLearning TIER_1 English(EN) · /u/Future_Caregiver_643 ·

    我构建了一个开源知识图谱管道,采用混合检索来改进 LLM 的多跳推理能力 [P]

    <!-- SC_OFF --><div class="md"><p>Hey everyone,</p> <p>I built an open-source full-stack pipeline (Django + React) that constructs a Knowledge Graph from raw text, detects thematic communities, and uses hybrid search to solve the &quot;lost in the middle&quot; problem in standard…

  412. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    链式思考的“副作用”:三种推理适得其反的任务

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GX38N645" rel="noopener noreferrer">Prompt Engineering Pocket Guide: Techniques for Getting the Most from LLMs</a> </li> <li> <strong>Also by me:</strong> <em>Thinking in Go</em> (2-book series) — <a href="http…

  413. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    人工智能在数学推理中的应用:语言模型、神经符号系统与验证发现的综合性调查

    "Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery" We critically assess failure modes: brittleness under perturbation, reward hacking, multimodal grounding failures, fragile formalization, …

  414. dev.to — LLM tag TIER_1 English(EN) · Alex Towell ·

    Value Functions Over Reasoning Traces

    <p>In <a href="https://metafunctor.com/post/2024-10-15-latent-reasoning-traces/" rel="noopener noreferrer">Latent Reasoning Traces</a>, I described a simple system: store successful reasoning traces, retrieve similar ones, use them to scaffold new problems. The traces serve as le…

  415. dev.to — LLM tag TIER_1 English(EN) · Alex Towell ·

    MCTS-Reasoning: LLM推理的树搜索

    <p>I've been working on applying Monte Carlo Tree Search to LLM reasoning. The idea: multi-step reasoning is a sequential decision problem, and MCTS is good at those.</p> <h2> The Problem with Single-Shot Reasoning </h2> <p>When you ask an LLM a hard question, it generates one re…

  416. dev.to — LLM tag TIER_1 English(EN) · Alex Towell ·

    潜在推理痕迹:记忆作为学习到的先验

    <p>Every time you ask an LLM a question, it reasons from scratch. All that computation (the chain of thought, the intermediate steps, the successful pattern that led to a correct answer) evaporates the moment the response is complete.</p> <p>The model doesn't learn from its own s…

  417. dev.to — LLM tag TIER_1 English(EN) · keeper ·

    Gemma 4 12B:隐藏的推理成本

    <h1> Gemma 4 12B: The Hidden Reasoning Tax </h1> <h2> Motivation </h2> <p>I recently acquired an RTX 5060 Ti 16GB for local LLM inference and wanted to find the best model for my use case: technical writing, code generation, and analysis in Chinese. Google's Gemma 4 12B seemed li…

  418. dev.to — LLM tag TIER_1 English(EN) · pixelbank dev ·

    知识蒸馏 — 深度解析 + 问题:罗马数字转整数

    <p><em>A daily deep dive into llm topics, coding problems, and platform features from <a href="https://pixelbank.dev" rel="noopener noreferrer">PixelBank</a>.</em></p> <h2> Topic Deep Dive: Knowledge Distillation </h2> <p><em>From the Deployment &amp; Optimization chapter</em></p…

  419. r/MachineLearning TIER_1 English(EN) · /u/zdeneklapes ·

    使用监督学习还是强化学习微调推理LLM?[D]

    <!-- SC_OFF --><div class="md"><p>Hello,</p> <p>I have a task to fine-tune small LLMs on annotated conversational data. The dataset contains not only the final answers, but also reasoning traces and tool-calling decisions (i.e., when the model should think and when it should call…

  420. dev.to — LLM tag TIER_1 English(EN) · Алексей Гормен ·

    你的 AI 有两个大脑:快速模式和 A11 深度推理引擎

    <p>In most tasks, a system relies on <strong>high‑speed thinking driven by attention vectors</strong> this is <em>intuition</em>.<br /><br /> It is a <strong>fast, energy‑efficient, pattern‑oriented mode</strong>, which can be described as:</p> <p><strong>Fast Pattern Heuristics …

  421. r/MachineLearning TIER_1 English(EN) · /u/Sensitive_Air_5745 ·

    冗长不等于忠实:一个关于推理模型无法执行忠实推理的架构论证 [D]

    <!-- SC_OFF --><div class="md"><p>Essay argues that reasoning models cannot perform faithful inference because their reasoning trace and final answer come from the same operation. Engages with Lanham/Turpin/Mirzadeh in empirical critique, and with HRM, TRM, GRAM, AlphaProof, and …