PulseAugur
实时 11:02:04
English(EN) LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning

新研究聚焦大语言模型推理、长上下文和工具集成

多篇研究论文探讨了大语言模型(LLM)推理能力的进步,重点在于提高长时任务和工具集成的性能。Apple的研究引入了LEAD,一种通过整合未来验证和重叠回滚来克服LLM推理中“无恢复瓶颈”的方法,使模型能够解决跳棋跳跃等复杂问题。其他论文提出了用于工具集成推理中的逐轮回溯自我蒸馏的TurnSight,衡量解释器稳定性的方法,以及通过测试时缩放和从失败轨迹中反思性学习来分析和改进推理过程的技术。此外,研究还探讨了优化推理接口以缩短响应时间,并通过新颖的记忆机制评估长上下文推理。 AI

影响 这些进展旨在提高大语言模型在复杂推理任务中的可靠性、效率和能力,可能催生更复杂的AI应用。

排序理由 多篇在arXiv上发表的学术论文,详细介绍了用于大语言模型推理的新方法和分析。

在 Apple Machine Learning Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 228 个来源。 我们如何撰写摘要 →

新研究聚焦大语言模型推理、长上下文和工具集成

报道来源 [228]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    LEAD:打破长时推理中的“无恢复”瓶颈

    Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposition is essential for stability, extreme decomposition creates a “no-recovery bot…

  2. arXiv cs.CL TIER_1 English(EN) · Bingxuan Li, Simo Du, Yue Guo ·

    面向自学习诊断代理的推理与双记忆联合优化

    arXiv:2604.07269v2 Announce Type: replace Abstract: Clinical expertise improves not only by acquiring medical knowledge, but by accumulating experience that yields reusable diagnostic patterns. Recent LLMs-based diagnostic agents have shown promising progress in clinical reasonin…

  3. arXiv cs.CL TIER_1 English(EN) · Hamed Damirchi, Ignacio Meza De la Jara, Damith Ranasinghe, Yuhang Liu, Javen Shi ·

    LLM残差流轨迹中的推理错误具有区域和方向

    arXiv:2608.05660v1 Announce Type: cross Abstract: As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. Recent trajectory-based methods seek th…

  4. arXiv cs.CL TIER_1 English(EN) · Xinye Wang, Junxiao Liu, Shujian Huang ·

    RP-OPSD:面向多语言推理迁移的推理-枢轴引导的在线策略自蒸馏

    arXiv:2608.06347v1 Announce Type: new Abstract: Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising paradigm…

  5. arXiv cs.CL TIER_1 English(EN) · Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han ·

    多语言数学推理的策略内Delta蒸馏

    arXiv:2608.05802v1 Announce Type: new Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Pol…

  6. arXiv cs.CL TIER_1 English(EN) · Hongbo Ma, Bangji Yang, Yunqian Selina Cheng, Jiajun Fan, Hanwen Zhang, Ge Liu ·

    Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving

    arXiv:2608.05254v1 Announce Type: new Abstract: Large language models can derive a plausible mathematical object yet still violate explicit requirements--for example, by omitting a modular reduction, returning a non-integer, or using the wrong encoded answer form. We introduce Co…

  7. arXiv cs.CL TIER_1 English(EN) · Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs ·

    LLM中的条件认知偏差:有偏见的对话轮次如何调节上下文内推理

    arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles t…

  8. arXiv cs.AI TIER_1 English(EN) · Emanuele Marconato, Samuele Bortolotti, Emile van Krieken, Paolo Morettin, Elena Umili, Antonio Vergari, Efthymia Tsamoura, Andrea Passerini, Stefano Teso ·

    神经符号AI中的符号接地:推理捷径的温和介绍

    arXiv:2510.14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or structural constraints. As such, it represents one of the most promising avenues for reliable a…

  9. arXiv cs.AI TIER_1 English(EN) · James Xu Zhao, Bryan Hooi, See-Kiong Ng ·

    面向知识密集型任务,推理模型的测试时缩放目前仍无效

    arXiv:2509.06861v3 Announce Type: replace Abstract: Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. However, frontier models still suffer from factuality hallucinations, raising the…

  10. arXiv cs.AI TIER_1 English(EN) · Hao Ai ·

    大型语言模型中思维链推理的均场动力学

    arXiv:2608.05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, w…

  11. arXiv cs.AI TIER_1 English(EN) · ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng ·

    DASH:用于推理模型在线策略自蒸馏的分歧自适应监督范围

    arXiv:2608.06243v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-…

  12. arXiv cs.AI TIER_1 English(EN) · Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen ·

    精炼优于重采样:LLM推理的测试时自纠正

    arXiv:2608.05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning dive…

  13. arXiv cs.AI TIER_1 English(EN) · Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Yang Li, Deguo Xia, Jizhou Huang ·

    啄木鸟蒸馏:弱模型诊断强模型中的推理错误

    arXiv:2608.05168v1 Announce Type: new Abstract: Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise from localized reasoning bugs in intermediate steps rather than from global incompetence. We…

  14. arXiv cs.AI TIER_1 English(EN) · Yu Gu, Zhi Zheng, Yunpeng Ba, Xialiang Tong, Mingxuan Yuan, Zhenkun Wang ·

    Hyper-ES:通过下降方向合并实现LLM推理的有效进化策略

    arXiv:2608.05541v1 Announce Type: new Abstract: Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such h…

  15. arXiv cs.LG TIER_1 English(EN) · Zonghuan Xu ·

    条件查询的假设检验:可学习性与交互的价值

    arXiv:2608.06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We f…

  16. arXiv cs.AI TIER_1 English(EN) · Byoungjae Min, Kennedy Edemacu, Sae-Hong Cho, Yoonhyuk Choi, Beakcheol Jang, Jong Wook Kim ·

    当缺失即是证据:评估大型语言模型中对完整性敏感的负面推理

    arXiv:2608.04591v1 Announce Type: cross Abstract: Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, th…

  17. arXiv cs.AI TIER_1 English(EN) · Qiyuan Zhu, Dezhi Li, Pengyu Cheng, Tianle Chen, Jiacheng Wang, Ruijie Shen, Hao Gu, Sida Lin, Zirui Liu, Jiacheng Liu, Sirui Han ·

    更少Token,更小缓存:奖励协调的高效推理

    arXiv:2608.04771v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inference cost. KV-cache compression is a common soluti…

  18. arXiv cs.CL TIER_1 English(EN) · Damien Sileo, Valentin Lacombe, Dimitri Kachler ·

    Reasoning Core: 为完成监督推理训练设计广泛的程序化数据

    arXiv:2608.05148v1 Announce Type: new Abstract: Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tuning. We introduce Reasoning Core, a collection of 50 generators spanning mathema…

  19. arXiv cs.CL TIER_1 English(EN) · Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora ·

    迈向技能原生大模型:利用技能熵进行长时推理的基准测试和训练

    arXiv:2608.05139v1 Announce Type: new Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill…

  20. arXiv cs.CL TIER_1 English(EN) · Ian B. de Haan, Peter van der Putten, Max van Duijn ·

    评估推理模型中的心智理论:鲁棒性优于推理

    arXiv:2608.04646v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via …

  21. arXiv cs.CL TIER_1 English(EN) · Bhiman Kumar Baghel, Anna Chrabaszcz, Tessa Warren, Michael Walsh Dickey, Haley C. Dresang, Xiang Lorraine Li ·

    STRIVE:探究分级合理性生成与评估中的推理极限

    arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled eve…

  22. arXiv cs.AI TIER_1 English(EN) · Junqi Chen, Sirui Chen, Chaochao Lu ·

    训练后能否将大型语言模型转变为因果推理模型?

    arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in this domain, their precise causal estimation capabilities are still limited, and…

  23. arXiv cs.AI TIER_1 English(EN) · Purbesh Mitra, Sennur Ulukus ·

    用于多轮推理的链式递归语言模型

    arXiv:2608.05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the fin…

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    ChronoVision:通过潜在状态重建实现时间推理

    Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulat…

  25. Hugging Face Daily Papers TIER_1 English(EN) ·

    多语言数学推理的策略内Delta蒸馏

    On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical…

  26. Hugging Face Daily Papers TIER_1 English(EN) ·

    迈向技能原生大模型:利用技能熵进行长时推理的基准测试和训练

    Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose step…

  27. arXiv cs.AI TIER_1 English(EN) · Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary ·

    推理大模型中的测试时缩放:推理模式、评估与可复现性

    arXiv:2608.04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single traje…

  28. arXiv cs.LG TIER_1 English(EN) · Jian Zhang, Bingyi Wang, Yizhi Liu ·

    CausalOPD:用于蒸馏因果链推理的首个错误步骤监督

    arXiv:2608.03673v1 Announce Type: new Abstract: Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning. Al…

  29. arXiv cs.CL TIER_1 English(EN) · Ming Shen, Zhikun Xu, Jacob Dineen, Xiao Ye, Ben Zhou ·

    BOW:训练语言模型推理合理的下一个词

    arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word reasoning methods make this tension explicit: the…

  30. arXiv cs.AI TIER_1 English(EN) · Daeyoung Roh, Donghee Han ·

    推理失败之前:Agentic RAG中的证据前程序性失败

    arXiv:2608.02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a proced…

  31. arXiv cs.AI TIER_1 English(EN) · Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang, Chen Zhang, Shunian Chen, Caijun Yan, Jianyao Xu, Shunyu Liu, Weijie Fu, Peiliang Li, Xiaozhi Chen, Yuxiang Cai ·

    SpatialCLI:学习使用空间工具进行推理,然后不使用它们

    arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mism…

  32. arXiv cs.AI TIER_1 English(EN) · Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu ·

    TurnSight:面向工具集成推理的逐轮事后自蒸馏

    arXiv:2608.04007v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit ass…

  33. arXiv cs.AI TIER_1 English(EN) · Francesca Carlon, Vincent Ginis, Andres Algaba ·

    更短的推理,更早的答案?对推理界面的评估

    arXiv:2608.03401v1 Announce Type: cross Abstract: Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluat…

  34. arXiv cs.AI TIER_1 English(EN) · Shashwat Sourav, Aishwarya Balwani ·

    The Tell-Tale Trace: 使用思维链动力学检测大型语言模型中的推理失败

    arXiv:2608.03291v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning…

  35. arXiv cs.AI TIER_1 English(EN) · Simon Lam-Muir ·

    点火真实存在,并体现在读出中:循环深度推理器中的潜在构成、难度时钟点火以及接口构成提交

    arXiv:2608.03263v1 Announce Type: cross Abstract: We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal training data. We grow an independent realization of a published 30M-parameter …

  36. arXiv cs.AI TIER_1 English(EN) · Dawei Liu, Haixu Song, Shuang Cheng, Shijie Wang, Haozheng Hou, Kaifeng Liu, Ermo Hua, Zhonghang Yuan, Zhijie Zhong, Yuchen Fan, Biqing Qi, Bowen Zhou ·

    PI-Mem:将长上下文推理能力推至 3.6M tokens,采用并行迭代记忆

    arXiv:2608.03048v1 Announce Type: cross Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irr…

  37. arXiv cs.AI TIER_1 English(EN) · Eddie Conti, \'Alvaro Parafita, Axel Brando ·

    通过归因可分离性衡量解释器稳定性

    arXiv:2608.02697v1 Announce Type: cross Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definitio…

  38. arXiv cs.AI TIER_1 English(EN) · Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua ·

    ReflectRL:通过反思到直接推理从黄金负面轨迹中学习

    arXiv:2608.03972v1 Announce Type: new Abstract: On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the exper…

  39. Hugging Face Daily Papers TIER_1 English(EN) ·

    迈向技能原生大模型:利用技能熵进行长时推理的基准测试和训练

    Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose step…

  40. Hugging Face Daily Papers TIER_1 English(EN) ·

    ReflectRL:通过反思到直接推理从黄金负轨迹中学习

    On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-…

  41. Hugging Face Daily Papers TIER_1 English(EN) ·

    CausalOPD:用于蒸馏因果链推理的首个错误步骤监督

    Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning. Although large language models perform well on suc…

  42. Hugging Face Daily Papers TIER_1 English(EN) ·

    当多个答案都有效时,投票机制失效:LLM中K选最佳因果推理的符号验证

    Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a…

  43. arXiv cs.CL TIER_1 English(EN) · Sean Gip Lim, William Chandra Tjhi, Hai Leong Chieu ·

    低资源东南亚语言的原生多语言思维链推理

    arXiv:2608.00533v1 Announce Type: new Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex lo…

  44. arXiv cs.CL TIER_1 English(EN) · Kang Liu, Zijing Wang, Yongkang Liu, Mengjie Zhao, Xiaocui Yang, Shi Feng, Yifei Zhang, Daling Wang ·

    TRAM:利用轨迹推导的辅助记忆增强多模态推理

    arXiv:2608.01922v1 Announce Type: new Abstract: Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as reasoning trajectories grow, models may become less effective at using informat…

  45. arXiv cs.LG TIER_1 English(EN) · Hector Zenil, Luan Ozelim ·

    以精确贝叶斯最优标准衡量语言模型中的上下文算法推理

    arXiv:2608.01575v1 Announce Type: new Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning …

  46. arXiv cs.CL TIER_1 English(EN) · Zhaoxin Yu, Qi Shen, Hengli Li, Zhaowei Zhang, Song-Chun Zhu, Chi Zhang, Zilong Zheng ·

    GradCuit:信用分配梯度流实现鲁棒且可解释的测试时潜在推理

    arXiv:2608.02585v1 Announce Type: cross Abstract: Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these sta…

  47. arXiv cs.CL TIER_1 English(EN) · Mengting Ai, Jingrui He, Yue Guo ·

    准确性等于证据吗?KV缓存压缩下的推理忠实度

    arXiv:2608.01631v1 Announce Type: new Abstract: KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that supports it. We test this assumption for large reasoning models and show that it c…

  48. arXiv cs.CL TIER_1 English(EN) · Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao ·

    答案正确方法错误:捷径式攻击误导了LLM在前沿科学基准上的推理评估

    arXiv:2608.02442v1 Announce Type: cross Abstract: Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify …

  49. arXiv cs.CL TIER_1 English(EN) · Shigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan, Anbang Yao ·

    关注自身思想:通过1.58位量化视角打破推理LLM训练后量化的瓶颈

    arXiv:2608.01078v1 Announce Type: new Abstract: We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to exhibit chain-of-thought reaso…

  50. Hugging Face Daily Papers TIER_1 English(EN) ·

    TurnSight:用于工具集成推理的逐轮事后自蒸馏

    Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy s…

  51. Hugging Face Daily Papers TIER_1 English(EN) ·

    ReflectRL:通过反思到直接推理从黄金负轨迹中学习

    On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-…

  52. Hugging Face Daily Papers TIER_1 English(EN) ·

    当答案众多且有效时,投票机制失效:LLM中K选最佳因果推理的符号验证

    Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a…

  53. Latent Space (swyx) TIER_1 English(EN) ·

    The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.

  54. arXiv cs.AI TIER_1 English(EN) · Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin ·

    关于链式思考忠实性的引导向量的泛化

    arXiv:2607.29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, model…

  55. arXiv cs.CL TIER_1 English(EN) · Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang, Kun Liang, Weijie Liu, Yunfang Wu ·

    端到端学习用于标量奖励模型的潜在推理轨迹

    arXiv:2607.29185v1 Announce Type: new Abstract: Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilistic reward modeling, they rely on superficial cues t…

  56. arXiv cs.AI TIER_1 English(EN) · Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng ·

    ThinkReset:可学习的中间接口构建用于有界上下文长时推理

    arXiv:2607.28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not…

  57. arXiv cs.AI TIER_1 English(EN) · Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley ·

    它思考的难度有多大?分析大型语言模型思维链轨迹中的步进式推理能耗

    arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into…

  58. arXiv cs.CL TIER_1 English(EN) · Sara Candussio, Daniel Scalena, Luca Bortolussi, Elisabetta Fersini, Malvina Nissim, Gabriele Sarti ·

    大型推理模型中基于熵的链式思考压缩选择方法解析

    arXiv:2607.28707v1 Announce Type: new Abstract: Entropy-based pruning has been proposed as an effective method for compressing Chain-of-Thought (CoT) reasoning with negligible accuracy loss. We test the robustness of low- and high-entropy CoT step selection methods across various…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    以精确贝叶斯最优标准衡量语言模型中的上下文算法推理

    Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning benchmark that supplies one exactly. Using the T…

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    GradCuit:信用分配梯度流实现鲁棒且可解释的测试时潜在推理

    Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded to…

  61. arXiv cs.CL TIER_1 English(EN) · Yecheng Wu, Song Han, Han Cai ·

    Lightning OPD 2.0:缓解大型推理模型交叉教师策略内蒸馏中的风格偏差

    arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the demonstrati…

  62. arXiv cs.LG TIER_1 English(EN) · Songshuo Lu, Zhi Chen, Yaohua Tang ·

    超越最佳教师:扩展和压缩推理解流形

    arXiv:2607.27770v1 Announce Type: new Abstract: A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore b…

  63. arXiv cs.CL TIER_1 English(EN) · Haojin Wang, Yike Wang, Shangbin Feng, Hannaneh Hajishirzi, Yulia Tsvetkov ·

    MentorCollab:用于高效推理的选择性大到小推理时引导

    arXiv:2602.05307v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve strong performance by producing long chains of thought, but their inference costs are high and often generate redundant reasoning. Small language models (SLMs) are far more efficient, yet st…

  64. arXiv cs.CL TIER_1 English(EN) · Zheng Wu, Chenhao Xue, Shijie Zheng, Yijie Lu, Cheng Yang, Zhuosheng Zhang ·

    你会走到洗车店吗?揭示大型语言模型在常识推理中的显著性偏差

    arXiv:2607.28478v1 Announce Type: new Abstract: As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a …

  65. arXiv cs.CL TIER_1 English(EN) · Amruta Parulekar, Jinu Lee, Dilek Hakkani-T\"ur, Hari Sundaram ·

    推理共识:通过加权有向无环图聚合实现LLM推理的结构集成

    arXiv:2607.27783v1 Announce Type: new Abstract: Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously …

  66. arXiv cs.AI TIER_1 English(EN) · Vivek Shukla, Varun Shukla, Atul, Divya Mishra, Mehul Kumar Das ·

    一种无参考评分方法用于检测大型语言模型中的隐式推理失败

    arXiv:2607.26102v1 Announce Type: cross Abstract: Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally …

  67. arXiv cs.CL TIER_1 English(EN) · Arnav Hiray, Agam Shah, Caleb Lu, Meghaj Tarte, Harsit Mittal, Sudheer Chava ·

    信用卡、困惑、计算与后果:我们能从语言模型推理中学到什么?

    arXiv:2607.26952v1 Announce Type: new Abstract: We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers nat…

  68. arXiv cs.CL TIER_1 English(EN) · Xuyao Feng, Anthony Hunter ·

    使隐含前提显性化,以实现对不完全三段论的逻辑理解

    arXiv:2603.06114v2 Announce Type: replace Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language processing (NLP) methods for handling enthymemes can potentially identify enthymemes i…

  69. arXiv cs.CL TIER_1 English(EN) · Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, Aishwarya Balwani ·

    探究推理性能的起源:强化学习与监督微调模型在数学问题解决中的表征质量

    arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage…

  70. Hugging Face Daily Papers TIER_1 English(EN) ·

    你会走到洗车行吗?揭示大型语言模型在常识推理中的显著性偏差

    As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bi…

  71. Hugging Face Daily Papers TIER_1 English(EN) ·

    信用卡、困惑、计算与后果:我们能从语言模型推理中学到什么?

    We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. W…

  72. arXiv cs.AI TIER_1 English(EN) · Jianshuo Dong, Yujia Fu, Chuanrui Hu, Chao Zhang, Han Qiu ·

    理解大型推理模型的认知习惯

    arXiv:2506.21571v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monitoring model behaviors. Inspired by the obse…

  73. arXiv cs.CL TIER_1 English(EN) · Zhichao Yan, Yunxiao Zhao, Jiapu Wang, Jiaoyan Chen, Xiaoli Li, Ru Li, Jeff Z. Pan ·

    超越事实准确性:使用 LogicScore 评估 RAG 系统中的全局推理完整性

    arXiv:2601.15050v5 Announce Type: replace Abstract: Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from \textit{factual myopia}: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation. This driv…

  74. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越自我认知:在大型语言模型中传播推理和检索中的不确定性

    Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language models can serve as an actionable signa…

  75. arXiv cs.AI TIER_1 English(EN) · Ahmed Haj Ahmed, Alvin Grissom II ·

    ADAGE:一种语言无关的类比推理评估流水线

    arXiv:2607.23058v1 Announce Type: cross Abstract: Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-…

  76. arXiv cs.AI TIER_1 English(EN) · Zirong Chen, Meiyi Ma ·

    Reason Popper-ly:用归纳逻辑编程修补上下文推理

    arXiv:2607.23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic …

  77. arXiv cs.LG TIER_1 English(EN) · Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto ·

    Re-FORC:用于高效思维链推理的自适应奖励预测

    arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on…

  78. arXiv cs.CL TIER_1 English(EN) · Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin ·

    LeAct:从专家行动中学习推理

    arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classic…

  79. arXiv cs.CL TIER_1 English(EN) · Xilun Chen, Ilia Kulikov, Vincent-Pierre Berges, Barlas O\u{g}uz, Rulin Shao, Gargi Ghosh, Jason Weston, Wen-tau Yih ·

    学习推理以提高事实准确性

    arXiv:2508.05618v2 Announce Type: replace Abstract: Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than their non-reasoning counterparts on long-form fac…

  80. Hugging Face Daily Papers TIER_1 English(EN) ·

    图表能否帮助大型语言模型进行推理?来自三段论推理的证据

    Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four repres…

  81. arXiv cs.AI TIER_1 English(EN) · Anthony Zhan ·

    用于扩散语言模型推理的简单策略梯度

    arXiv:2510.04019v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-training techniques, including reinforcement learning (RL), remains a key challenge for …

  82. arXiv cs.AI TIER_1 English(EN) · Kyle Cox, Darius Kianersi, Adri\`a Garriga-Alonso ·

    链式思考中的事后推理:解码和引导预承诺的答案

    arXiv:2603.01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisio…

  83. arXiv cs.AI TIER_1 English(EN) · Manqing Liu, David Williams-King, Ida Caspary, Linh Le, Hannes Whittingham, Puria Radmard, Cameron Tice, Edward James Young ·

    诊断推理模型中的病理性思维链

    arXiv:2602.13904v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT reasoning may exhibit failure modes that we note as pathologies, which prevent …

  84. arXiv cs.AI TIER_1 English(EN) · Renuka Oladri, Niveda Jawahar, Abdirisak Mohamed ·

    Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

    arXiv:2607.21433v1 Announce Type: cross Abstract: Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion (non-converged…

  85. arXiv cs.AI TIER_1 English(EN) · Adam Kostka (Warsaw University of Technology), Jaros{\l}aw A. Chudziak (Warsaw University of Technology) ·

    动态认知划分下的可解释信念协调

    arXiv:2607.21210v1 Announce Type: cross Abstract: Existing approaches to multi-agent belief combination have established mature foundations for combining uncertain beliefs under common assumptions: consensus methods use iterative averaging, logic-based methods resolve conflicting…

  86. arXiv cs.CL TIER_1 English(EN) · Ishan S. Kshirsagar ·

    沉默的重量:关于潜在国际象棋推理中权重而非草稿的因果论证

    arXiv:2607.20952v1 Announce Type: cross Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad the model actively consults during infer…

  87. arXiv cs.AI TIER_1 English(EN) · Wen Ye, Yuxiao Qu, Aviral Kumar, Xuezhe Ma ·

    MIRROR:从另一视角学习以进行多模态推理

    arXiv:2607.21552v1 Announce Type: new Abstract: Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text v…

  88. arXiv cs.AI TIER_1 English(EN) · Bartolomeo Bogliolo ·

    Euclid-MCP:用于通过Prolog实现确定性逻辑推理的模型上下文协议服务器

    arXiv:2607.21412v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approac…

  89. arXiv cs.AI TIER_1 English(EN) · Kilian Rueckschloss (Eberhard Karls Universitaet Tuebingen), Felix Weitkaemper (German University of Digital Science) ·

    规则如何表征因果知识:概率逻辑编程的因果建模

    arXiv:2607.21208v1 Announce Type: new Abstract: Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge supports only conclusions drawn from observations. His theory of causality, however, is developed …

  90. arXiv cs.AI TIER_1 English(EN) · Akihiro Takemura (National Institute of Informatics, Tokyo, Japan), Katsumi Inoue (National Institute of Informatics, Tokyo, Japan) ·

    可微逻辑编程用于缓解神经符号系统中的推理捷径

    arXiv:2607.21185v1 Announce Type: new Abstract: Neurosymbolic (NeSy) systems integrate neural networks with logical reasoning to achieve both generalization and interpretability, but recent work has shown they are susceptible to shortcut reasoning behaviors. We propose a novel me…

  91. arXiv cs.AI TIER_1 English(EN) · Sagnik Nath, Edith Aurora Graf, Liang Zhang, Diego Zapata-Rivera ·

    大型语言模型在数学问题解决中受可执行推理约束下的表征鲁棒性

    arXiv:2607.20520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failure…

  92. arXiv cs.AI TIER_1 English(EN) · Sizhe Tang, Guangyu Jiang, Yu Li, Rongqian Chen, Ioannis G. Kevrekidis, Tian Lan ·

    FlowEdit:信息论控制LLM推理流以解决涉及冲突的病态问题

    arXiv:2607.20500v1 Announce Type: new Abstract: Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in the open world can become ill-posed due to inconsistent conditions, conflicting statements, or …

  93. arXiv cs.CL TIER_1 English(EN) · Zhensheng Jin, Xin Dai, Zhenghao Liu, Chaojun Xiao, Huiyuan Xie, Yu Gu, Ge Yu, Maosong Sun ·

    REFACT:用于紧凑且忠实思维链推理的自适应事实重述

    arXiv:2607.20833v1 Announce Type: new Abstract: Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge. Existing…

  94. arXiv cs.CL TIER_1 English(EN) · Simone Angarano, Francesco Bertolotti, Federico D'Ambrosio, Michele Resta, Alessandro Rognoni, Nicol\`o Ruggeri, Dario Salvati, Andrea Valenti, Alberto Veneri, Martin Cimmino ·

    Domyn-Small:一个欧洲10B推理语言模型

    arXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an initial pre-training phase on 9 trillion tokens multilingual data, followed by a …

  95. Hugging Face Daily Papers TIER_1 English(EN) ·

    MIRROR:从另一视角学习以实现多模态推理

    Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit diff…

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    Euclid-MCP: 用于通过Prolog进行确定性逻辑推理的模型上下文协议服务器

    Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approaches address this gap by coupling neural models w…

  97. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jarosław A. Chudziak ·

    动态认知划分下的可解释信念协调

    Existing approaches to multi-agent belief combination have established mature foundations for combining uncertain beliefs under common assumptions: consensus methods use iterative averaging, logic-based methods resolve conflicting knowledge bases, and epistemic logic analyzes age…

  98. Hugging Face Daily Papers TIER_1 English(EN) ·

    规则如何表征因果知识:概率逻辑编程的因果建模

    Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge supports only conclusions drawn from observations. His theory of causality, however, is developed exclusively within Bayesian networks and causal …

  99. arXiv cs.AI TIER_1 English(EN) · Xinbang Dai, Zheyu Xin, Huikang Hu, Lin Ren, Rihui Jin, Guohui Xiao, Guilin Qi, Kuicai Dong, Zhaocheng Du, Yuyang Zhang ·

    EvoThink:通过自剪枝和顿悟偏好优化实现大型推理模型的思维演进

    arXiv:2607.19962v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switching and reasoning trajectory compression, fail to ma…

  100. arXiv cs.AI TIER_1 English(EN) · Yousef Khan, Luca Gherardini, Marco Maratea, Joel Arrais, Jose Sousa ·

    CLARK: 用于知识图谱自适应推理的闭环学习

    arXiv:2607.19996v1 Announce Type: new Abstract: Machine Learning models are widely used for automating classification tasks by extracting statistical patterns from data. However, their performance deteriorates if the data distribution changes, making them ill-suited to handle unc…

  101. arXiv cs.AI TIER_1 English(EN) · Alexis Fox, Junlin Wang, Paul Rosu, Bhuwan Dhingra ·

    PRO-LONG:程序化记忆实现长时域推理

    arXiv:2607.20064v1 Announce Type: new Abstract: Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks s…

  102. arXiv cs.AI TIER_1 English(EN) · Anmol Kankariya, Sercan \"O. Ar{\i}k ·

    PoTRE:受认知异质性启发的测试时推理

    arXiv:2607.20268v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle…

  103. arXiv cs.AI TIER_1 English(EN) · Hiskias Dingeto ·

    训练模型而非读者:用于可验证激活解释的解码监督

    arXiv:2607.20379v1 Announce Type: new Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claim…

  104. arXiv cs.AI TIER_1 English(EN) · Wael AbdAlmageed ·

    SoftReason:一种在த்தனை感知数据上实现全可微的神经-软符号演绎推理架构

    arXiv:2607.20402v1 Announce Type: new Abstract: In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledg…

  105. arXiv cs.AI TIER_1 English(EN) · Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean, Maria Liakata ·

    开放式问答中推理的无参考评估

    arXiv:2607.19678v1 Announce Type: cross Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for …

  106. arXiv cs.AI TIER_1 English(EN) · Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li ·

    SLPO:通过代理策略扩展潜在推理

    arXiv:2607.19691v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediat…

  107. arXiv cs.AI TIER_1 English(EN) · Miao Li, Alexander Gurung, Irina Saparina, Mirella Lapata ·

    SciTrek:评估和改进科学文献中的长上下文数值推理

    arXiv:2509.21028v4 Announce Type: replace Abstract: We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs). Existing long-context datasets with inputs beyond 64K tokens either targe…

  108. arXiv cs.AI TIER_1 English(EN) · Wouter W. L. Nuijten, Bert de Vries ·

    来自认知先验的复杂策略

    arXiv:2607.19518v1 Announce Type: new Abstract: Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its central computational role is simpler: within a planning horizon, it makes active inference …

  109. Hugging Face Daily Papers TIER_1 English(EN) ·

    REFACT:用于紧凑且忠实思维链推理的自适应事实重述

    Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge. Existing grounding methods either attach citations after…

  110. arXiv cs.CL TIER_1 English(EN) · Polina Tsvilodub, Fausto Carcassi, Michael Franke ·

    具有灵活生成意义和表达替代方案的语用推理计算模型

    arXiv:2607.18443v1 Announce Type: new Abstract: Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must…

  111. arXiv cs.CL TIER_1 English(EN) · Abir Harrasse, Michael Lan, Hunar Batra, Fateme Hashemi Chaleshtori, Chaithanya Bandi ·

    推理微调诱导持久的潜在策略状态

    arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improv…

  112. arXiv cs.AI TIER_1 English(EN) · Liam Swayne ·

    Relay-Bench:在多领域推理链上评估LLM

    arXiv:2607.18438v1 Announce Type: cross Abstract: Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The …

  113. arXiv cs.AI TIER_1 English(EN) · Ayhan Suleymanzade, Halil Alperen Gozeten, Michael Bronstein, \.Ismail \.Ilkan Ceylan, Jinwoo Kim ·

    MUX:通过多路复用令牌实现连续推理

    arXiv:2607.18264v1 Announce Type: new Abstract: Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoning step conveys only a single subword, and many are …

  114. arXiv cs.AI TIER_1 English(EN) · Dmitrii Kharlapenko, Terry Jingchen Zhang, Arth Singh, Alessandro Stolfo, Arthur Conmy, Mrinmaya Sachan, Zhijing Jin ·

    流式推理表征

    arXiv:2602.04843v2 Announce Type: replace Abstract: Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking. Yet we lack a mechanistic account of how extended thinking changes hidden-state representations ove…

  115. arXiv cs.AI TIER_1 English(EN) · Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang ·

    少复制,多接地:通过证据感知强化学习克服长上下文推理中的重复复制问题

    arXiv:2607.19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical…

  116. arXiv cs.AI TIER_1 English(EN) · Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi ·

    离境GRPO:利用特权信息学习解决难题的推理能力

    arXiv:2607.19313v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{ze…

  117. arXiv cs.AI TIER_1 English(EN) · Yu Wang, Ming Fan, Xicheng Zhang, Zhiyong Li, Zhihu Wang, Caiyue Xu, Dahai Hu, Ting Liu ·

    DAIS: 依赖感知中间问答监督用于复杂推理

    arXiv:2607.19088v1 Announce Type: cross Abstract: Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. W…

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    训练模型而非读者:用于可验证激活解释的可解码性监督

    Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the recon…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    SLPO:通过代理策略扩展潜在推理

    Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent…

  120. Hugging Face Daily Papers TIER_1 English(EN) ·

    DAIS: 依赖感知中间问答监督用于复杂推理

    Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Super…

  121. arXiv cs.AI TIER_1 English(EN) · Satyam Kumar, Saurabh Jha ·

    CADENCE:通过覆盖自适应策略内蒸馏缩小推理差距

    arXiv:2607.16955v1 Announce Type: cross Abstract: On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to…

  122. arXiv cs.AI TIER_1 English(EN) · Thorir Mar Ingolfsson, Wajeeha Tahir, Anna Tegon, Lionnus Kesting, Gamze \.Islamo\u{g}lu, Luca Benini ·

    递归推理模型的量化

    arXiv:2607.16237v1 Announce Type: cross Abstract: Recursive reasoning models solve hard puzzles by applying compact, weight-tied blocks over many refinement steps. Because these blocks are reused many times, quantizing them creates a unique dynamical problem: the quantization err…

  123. arXiv cs.AI TIER_1 English(EN) · Brian K Chen ·

    压力下的逻辑判断:用学习到的软前缀诊断三段论稳定性

    arXiv:2607.18228v1 Announce Type: new Abstract: To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we chara…

  124. arXiv cs.AI TIER_1 English(EN) · Sheldon Yu, Tong Yu, Xunyi Jiang, Rohan Surana, Gagan Mundada, Sungchul Kim, Lina Yao, Julian McAuley, Junda Wu ·

    我们能否打破大型语言模型的自我循环?通过激活引导实现细粒度推理控制

    arXiv:2607.18100v1 Announce Type: new Abstract: Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches …

  125. arXiv cs.AI TIER_1 English(EN) · Zehua Cheng, Wei Dai, Jiahao Sun ·

    约束锚定推理轨迹

    arXiv:2607.16727v1 Announce Type: new Abstract: Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning. We find that in state-of-the-art open-sour…

  126. arXiv cs.AI TIER_1 English(EN) · Yanxiao Zhao, Yaqian Li, Zihao Bo, Rinyoichi Takezoe, Haojia Hui, Mo Guang, Lei Ren, Xiaolin Qin, Kaiwen Long ·

    SATQuest:用于LLM逻辑推理评估和强化微调的验证器

    arXiv:2509.00930v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit strong general reasoning, yet the community lacks controllable, scalable, and verifiable tools to analyze and improve these abilities. We present SATQuest, a verifier that generates diverse S…

  127. arXiv cs.AI TIER_1 English(EN) · Yosub Shin, Michael Buriek, Boris Sobolev, Pavel Bushuyeu, Vikas Kumar, Haoyang Xu, Samuel Watson, Igor Molybog ·

    固定训练协议下的多模态推理数据高效策展

    arXiv:2601.10922v2 Announce Type: replace Abstract: We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constant and the main degree of freedom is the training…

  128. Hugging Face Daily Papers TIER_1 English(EN) ·

    计算模型用于语用推理,并灵活生成意义和表达的替代方案

    Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that…

  129. arXiv cs.CL TIER_1 English(EN) · Leichao Dong, Dongxu Zhang, Yiding Sun, Qirui Wang, Yuhan Wang, Lin Chen, Jihua Zhu ·

    更好的开始,更好的结束:自举迭代自推理蒸馏用于压缩推理

    arXiv:2607.15736v1 Announce Type: new Abstract: Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations, repeated self-verification, and detours that do not improve the final answer. Ex…

  130. arXiv cs.AI TIER_1 English(EN) · Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov ·

    从预训练到后训练的推理理解

    arXiv:2607.16097v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basi…

  131. arXiv cs.AI TIER_1 English(EN) · Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur ·

    精确但脱节:评审员的精确度不保证多智能体数学推理中的批评被采纳

    arXiv:2607.15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 ver…

  132. arXiv cs.AI TIER_1 Italiano(IT) · J. N. Hooker ·

    逻辑、优化与人工智能

    arXiv:2607.15532v1 Announce Type: new Abstract: Logic and optimization can, in combination, make valuable contributions to rule-based AI. Logic is the obvious medium for encoding a rule base and drawing inferences from it, while optimization provides a powerful technology for com…

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    揭示语言模型中的潜在推理策略

    A language model p_θ(y mid x) trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution. We study the problem of decomposing the response distribution of a given p…

  134. arXiv cs.CL TIER_1 English(EN) · Seungone Kim, Ian Wu, Jinu Lee, Xiang Yue, Seongyun Lee, Mingyeong Moon, Carolin Lawrence, Kiril Gashteovski, Julia Hockenmaier, Graham Neubig, Sean Welleck ·

    使用推理模型作为评估器来扩展评估时计算

    arXiv:2503.19877v3 Announce Type: replace Abstract: As language model (LM) outputs get more and more natural, it is becoming more difficult than ever to evaluate their quality. Simultaneously, increasing LMs' "thinking" time through scaling test-time compute has proven an effecti…

  135. arXiv cs.AI TIER_1 English(EN) · Haohua Niu, Xingtong Yu, Yang Liu, Junfeng Fang, Xuanting Xie, Jie Tan, Zhongjian Zhang, Hong Cheng, Yuan Fang ·

    CoEvoT:图大语言模型推理的共进化思维链提示

    arXiv:2607.14114v1 Announce Type: cross Abstract: Graph learning under distribution shift presents a persistent challenge, where models adapt to new graphs with limited or even no supervision. Recent graph--LLM approaches move toward label-efficient prediction by linearizing grap…

  136. arXiv cs.AI TIER_1 English(EN) · Abdullah Shaikh, Zain Naqi, Taha Zahid, Sandesh Kumar, Abdul Samad ·

    HABIB_TAZ 在 SemEval-2026 Task 11:通过合成训练和多目标优化解耦形式逻辑与内容

    arXiv:2607.14349v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility. In this paper, we p…

  137. arXiv cs.AI TIER_1 English(EN) · Jungseob Lee, Seungyoon Lee, Suhyune Son, Dongyub Jude Lee, Sungbin Han, Sugyeong Eo, Heuiseok Lim ·

    Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

    arXiv:2607.14552v1 Announce Type: cross Abstract: A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fa…

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    从预训练到后训练的推理理解

    Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining ch…

  139. arXiv cs.CL TIER_1 English(EN) · Heuiseok Lim ·

    Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

    A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix shows the generator the gold ans…

  140. arXiv cs.AI TIER_1 English(EN) · Hefeng Zhou, Jinxuan Zhang, Jiong Lou, Yuxin Liu, Chaochao Lu, Jingjing Qu, Jie Li ·

    深度交互:大型推理模型的高效人机交互方法

    arXiv:2607.14049v1 Announce Type: new Abstract: The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically invol…

  141. arXiv cs.CL TIER_1 English(EN) · Abdul Samad ·

    HABIB_TAZ 在 SemEval-2026 Task 11:通过合成训练和多目标优化解耦形式逻辑与内容

    While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility. In this paper, we present our system for SemEval-2026 Task 11, which …

  142. arXiv cs.AI TIER_1 English(EN) · Jie Li ·

    深度交互:大型推理模型的高效人机交互方法

    The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically involve re-generating another response that may make …

  143. Hugging Face Daily Papers TIER_1 English(EN) ·

    闭环知识动力学:饱和与逃逸的操作框架

    Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback. We study why closed-loop knowledge systems saturate and what external information can m…

  144. arXiv cs.AI TIER_1 English(EN) · Junyu Ren ·

    基于证据的已验证代理推理:通过工具认证的内核证明消除经验推理中 LLM 幻觉的途径

    arXiv:2607.12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Ver…

  145. arXiv cs.CL TIER_1 English(EN) · Xinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan, Xiaochong Lan, Yifan Li, Yuchen Yan, Han Peng, Zican Dong, Zhenduo Zhang, Tianshu Wang, Xinyu Kong, Zujie Wen, Wayne Xin Zhao, Zhiqiang Zhang, Jun Zhou ·

    Ring-Zero:将零样本强化学习扩展到万亿参数以实现涌现式推理

    arXiv:2607.12395v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, exist…

  146. arXiv cs.AI TIER_1 English(EN) · Ning Liu ·

    追踪、排名、破解:认知工作记忆提升语言代理的多跳推理能力

    arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed,…

  147. arXiv cs.AI TIER_1 English(EN) · Dongyue Li, Zhenshuo Zhang, Minxuan Duan, Edgar Dobriban, Hongyang R. Zhang ·

    高效学习分支网络以进行多任务算法推理

    arXiv:2512.01113v2 Announce Type: replace-cross Abstract: Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reasoning abilities, designed for graph neural networks and also for transformer models. …

  148. arXiv cs.AI TIER_1 English(EN) · Julius Steiglechner, Lucas Mahler, Gabriele Lohmann ·

    大型语言模型能看到烟雾但看不到火:用 Elenchos 评估溯因推理

    arXiv:2607.12733v1 Announce Type: new Abstract: Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce E…

  149. Hugging Face Daily Papers TIER_1 English(EN) ·

    GSM-Plus-BN:面向大型语言模型的孟加拉语数学推理的基于扰动的基准测试

    The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Banglades…

  150. arXiv cs.CL TIER_1 English(EN) · Swastika Kundu ·

    GSM-Plus-BN:面向大型语言模型的孟加拉语数学推理的基于扰动的基准测试

    The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Banglades…

  151. arXiv cs.LG TIER_1 English(EN) · Gabriele Lohmann ·

    大型语言模型能看到烟但看不到火:用 Elenchos 评估溯因推理

    Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cros…

  152. arXiv cs.AI TIER_1 English(EN) · Junyu Ren ·

    基于证据的已验证代理推理:通过工具证明的核心证据消除LLM在经验推理中的幻觉的途径

    Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Agentic Reasoning), a Lean 4-based tool-call…

  153. arXiv cs.CL TIER_1 English(EN) · Jun Zhou ·

    Ring-Zero:将零样本强化学习扩展至万亿参数以实现涌现式推理

    Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small mode…

  154. arXiv cs.AI TIER_1 English(EN) · Sabari Iyyappan Duraipandian, Shreya Sanjay Boyane, Manju Nagesh, Jerome Francis, Archana Vaidheeswaran, Kevin Zhu ·

    将潜在的CoT推理解释为动力学系统

    arXiv:2607.09698v1 Announce Type: new Abstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate traces in the hidden space at each step, unlike explicit- CoT, which follows a sin…

  155. arXiv cs.AI TIER_1 English(EN) · Huan Zhu ·

    思考瓶颈:用于严谨归纳的沙漏推理

    arXiv:2607.11696v1 Announce Type: new Abstract: Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolat…

  156. arXiv cs.AI TIER_1 English(EN) · Joyjeet Singh ·

    平衡是初始化:物理结构深度平衡推理中的懒惰身份坍缩

    arXiv:2607.11116v1 Announce Type: cross Abstract: Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary …

  157. arXiv cs.AI TIER_1 English(EN) · Mohan Tang, Sidi Lu ·

    Turbo Connection:推理作为信息从高层到低层的流动

    arXiv:2602.17993v2 Announce Type: replace-cross Abstract: Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step informs the next. In this work, we adopt the perspective that the reasoning power of Tran…

  158. arXiv cs.AI TIER_1 English(EN) · Kejing Xia, Mingzhe Li, Lixuan Wei, Zhenbang Du, Xiangchi Yuan, Dachuan Shi, Qirui Jin, Wenke Lee ·

    MetaState:持久化工作记忆增强离散扩散语言模型的推理能力

    arXiv:2603.01331v3 Announce Type: replace-cross Abstract: Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each denoising step solely on the current hard-masked sequence, while intermediate cont…

  159. arXiv cs.AI TIER_1 English(EN) · Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi ·

    过滤推理得分:评估模型最自信轨迹的推理质量

    arXiv:2604.11996v2 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone does not reveal the quality of the reasoning used to produce it. This highlights a …

  160. arXiv cs.AI TIER_1 English(EN) · Mohammed Ehab, Aymane El Gadarri, Vivek F. Farias, Adam Jozefiak, Ciamac C. Moallemi ·

    OS-Pruner:通过最优停止法修剪推理模型的思维链

    arXiv:2607.11089v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning step…

  161. arXiv cs.AI TIER_1 English(EN) · JungMin Yun, JuneHyoung Kwon, YoungBin Kim ·

    CRiT-QA:使用反事实链和干扰陷阱评估多跳推理

    arXiv:2607.10562v1 Announce Type: new Abstract: Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks…

  162. arXiv cs.LG TIER_1 English(EN) · Ning Liu ·

    追踪、排名、破解:认知工作记忆提升语言智能体的多跳推理能力

    Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed, what it suspects, and what it still needs) lives …

  163. Hugging Face Daily Papers TIER_1 English(EN) ·

    Ring-Zero:将零样本强化学习扩展到万亿参数以实现涌现式推理

    Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small mode…

  164. arXiv cs.AI TIER_1 English(EN) · Huan Zhu ·

    思考瓶颈:沙漏推理用于严谨归纳

    Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolation between reasoning stages, so that informatio…

  165. Hugging Face Daily Papers TIER_1 English(EN) ·

    平衡即初始化:物理结构深度平衡推理中的惰性身份坍缩

    Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned ini…

  166. Hugging Face Daily Papers TIER_1 English(EN) ·

    OS-Pruner:通过最优停止法修剪推理模型的思维链

    Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improvi…

  167. Hugging Face Daily Papers TIER_1 English(EN) ·

    ProofCouncil:解决开放性数学问题的LLM智能体

    Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical ag…

  168. arXiv cs.AI TIER_1 English(EN) · Jakob Suchan, Julius Monsen, Salim Baloch, Mehul Bhatt ·

    Answer Set Programming 焕发活力!基于 ASP 和能量模型的端到端神经符号推理与学习

    arXiv:2607.08136v1 Announce Type: new Abstract: We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the con…

  169. arXiv cs.CL TIER_1 English(EN) · Chris Samarinas, Haw-Shiuan Chang, Hamed Zamani ·

    截断式步进采样结合过程奖励用于检索增强推理

    arXiv:2602.23440v4 Announce Type: replace Abstract: Reinforcement learning has emerged as an effective paradigm for training large language models to interleave reasoning with search engine calls. However, existing approaches face a fundamental credit assignment problem: methods …

  170. arXiv cs.AI TIER_1 English(EN) · Jack Hopkins, Dipika Khullar, Fabien Roger ·

    过度思考:放大推理权重以提取学习到的秘密

    arXiv:2607.08173v1 Announce Type: new Abstract: Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overt…

  171. arXiv cs.AI TIER_1 English(EN) · Shaoyu Wang, Kaiyue Zhao, Dongliang Wei, Przemys{\l}aw Andrzej Wa{\l}\k{e}ga, Dingmin Wang, Hongming Cai, Pan Hu ·

    Goal-Driven Reasoning in DatalogMTL with Magic Sets

    arXiv:2412.07259v5 Announce Type: replace Abstract: DatalogMTL is a powerful rule-based language for temporal reasoning. Due to its high expressive power and flexible modeling capabilities, it is suitable for a wide range of applications, including tasks from industrial and finan…

  172. arXiv cs.AI TIER_1 English(EN) · Fabien Roger ·

    过度思考:放大推理权重以提取学习到的秘密

    Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task ve…

  173. Hugging Face Daily Papers TIER_1 English(EN) ·

    Answer Set Programming 焕发活力!基于 ASP 和能量模型的端到端神经符号推理与学习

    We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the continuous latent space through explicit ASP-based …

  174. arXiv cs.CL TIER_1 English(EN) · Xinda Jia, Jinpeng Li, Zezhong Wang, Jingjing Li, Xingshan Zeng, Yasheng Wang, Weinan Zhang, Yong Yu, Weiwen Liu ·

    LLM 的快速、慢速和工具增强思维:一篇综述

    arXiv:2508.12265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable progress in reasoning across diverse domains. However, effective reasoning in real-world tasks requires adapting the reasoning strategy to the demands of the problem, ran…

  175. arXiv cs.AI TIER_1 English(EN) · Azwar Abdulsalam, Nishil Patel, Andrew Saxe ·

    RL Post-Training 构建组合推理策略

    arXiv:2607.07646v1 Announce Type: new Abstract: Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment wher…

  176. arXiv cs.AI TIER_1 English(EN) · Guangzhi Wang, Kai Li, Yinghao Jiao, Zhi Liu ·

    Refine Thought:用于嵌入模型推理的测试时推理方法

    arXiv:2511.13726v2 Announce Type: replace-cross Abstract: We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models. The method obtains the final semantic representation by running multiple forward passes of the text embeddi…

  177. arXiv cs.CL TIER_1 English(EN) · Ruilin Tong, Dong Gong ·

    MILES:用于自改进 LLM 推理的模块化指令记忆与可学习选择

    arXiv:2607.06974v1 Announce Type: new Abstract: Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience …

  178. arXiv cs.CL TIER_1 English(EN) · Hengyu Jin, Shu Yang, Di Wang ·

    最终检查点不足以分析训练轨迹中的潜在推理忠实度

    arXiv:2607.06648v1 Announce Type: cross Abstract: Latent reasoning methods perform multi-step inference entirely in the model's continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden states raise a question of faithfulness: whether …

  179. arXiv cs.CL TIER_1 English(EN) · Josip Juki\'c, Ivan Titov ·

    几何自蒸馏用于推理泛化

    arXiv:2607.06855v1 Announce Type: cross Abstract: On-policy distillation is a practical post-training recipe for large language models, supplying dense teacher supervision on the student's own trajectories. In privileged-context self-distillation, teacher and student are the same…

  180. arXiv cs.AI TIER_1 English(EN) · Dmitry Beresnev, Vladimir Makharev, Roman Khalikov, Ivan Oseledets, Petr Anokhin ·

    搜索、失败、恢复:面向可纠错推理的训练框架

    arXiv:2607.07492v1 Announce Type: new Abstract: Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent…

  181. arXiv cs.AI TIER_1 English(EN) · Andrew Saxe ·

    RL Post-Training 构建组合推理策略

    Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and ever…

  182. arXiv cs.AI TIER_1 English(EN) · Petr Anokhin ·

    搜索、失败、恢复:面向可纠错推理的训练框架

    Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent, a training and inference framework inspired by…

  183. arXiv cs.AI TIER_1 English(EN) · He Liu, Changtao Miao, Xinjie Yang, Tianle Song, Yin Wu, Junchi Chen, Bintao He, Xinyuan Zhang, Bo Zhang, Shi Yan, Wei Lu, Wei Wang, Danyang Xu, Jiansheng Cai, Zhe Li ·

    DT-Guard:面向无推理LLM安全护栏的意图驱动推理激活训练

    arXiv:2607.06326v1 Announce Type: new Abstract: Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. Existing guardrails face a practical trade-off betwee…

  184. arXiv cs.AI TIER_1 English(EN) · Yang Liu, Zhaokai Luo, Huayi Jin, Ruozhou He, Chenchen Hong, Zhiyong Wang, Yifei Liu, Yunfei Gu, Chentao Wu, Junhao Hu ·

    Akashic:一种具有 MemAttention 的低开销 LLM 推理服务

    arXiv:2607.05708v1 Announce Type: new Abstract: Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts in…

  185. arXiv cs.CL TIER_1 English(EN) · Dong Gong ·

    MILES:用于自改进LLM推理的可学习选择模块化指令记忆

    Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Exi…

  186. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agon:具有隐式竞争对手评分的竞争性跨模型强化学习

    Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for go…

  187. arXiv cs.CL TIER_1 English(EN) · Ivan Titov ·

    Geometric Self-Distillation for Reasoning Generalization

    On-policy distillation is a practical post-training recipe for large language models, supplying dense teacher supervision on the student's own trajectories. In privileged-context self-distillation, teacher and student are the same model conditioned on the same prefix, but the tea…

  188. arXiv cs.CL TIER_1 English(EN) · Di Wang ·

    最终检查点不足以分析训练轨迹中的潜在推理忠实度

    Latent reasoning methods perform multi-step inference entirely in the model's continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden states raise a question of faithfulness: whether these latent reasoning steps causally drive the fi…

  189. arXiv cs.AI TIER_1 English(EN) · Zhe Li ·

    DT-Guard:面向无推理LLM安全护栏的意图驱动推理激活训练

    Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. Existing guardrails face a practical trade-off between lightweight classification-based models, which…

  190. arXiv cs.AI TIER_1 English(EN) · Hehai Lin, Shilei Cao, Sudong Wang, Haotian Wu, Minzhi Li, Linyi Yang, Juepeng Zheng, Chengwei Qin ·

    LLM推理的交互式学习

    arXiv:2509.26306v5 Announce Type: replace Abstract: Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multiple Large Language Models (LLMs), thereby constructing stronger multi-agent systems (MAS). …

  191. arXiv cs.AI TIER_1 English(EN) · Jingchu Wang, Bingbing Xu, Yige Yuan, Dan Zhang, Bin Xie, Xiaoqian Sun, Huawei Shen ·

    R$^2$PO: 解耦 LLM 推理的部署和推理策略

    arXiv:2601.11960v3 Announce Type: replace-cross Abstract: Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincide with the one producing inference responses. We argue that this is a misleading in…

  192. arXiv cs.AI TIER_1 English(EN) · Mengqi Li, Lei Zhao, Anthony Man-Cho So, Ruoyu Sun, Xiao Li ·

    模型可自我提升:LLM推理的无奖励自训练

    arXiv:2510.18814v4 Announce Type: replace-cross Abstract: Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolving Post-Training (SePT), a simple post-tra…

  193. arXiv cs.CL TIER_1 English(EN) · Wei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao, Blake JianHang Chen, Ziqian Lin, Alec Go, Yu Meng ·

    深度思考,而非仅仅冗长:通过深度思考标记衡量大型语言模型的推理努力

    arXiv:2602.13517v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). However, recent findings suggest that raw token counts are unreliable proxies for rea…

  194. Hugging Face Daily Papers TIER_1 English(EN) ·

    Forethought:来自神经符号原始编程的可验证推理

    Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning …

  195. Hugging Face Daily Papers TIER_1 English(EN) ·

    RuleChef:将大语言模型任务知识锚定在人类可编辑规则中

    RuleChef utilizes large language models to generate and iteratively improve executable rules for NLP tasks through example-based learning and human feedback, resulting in fast and inspectable rule systems.

  196. arXiv cs.CV TIER_1 English(EN) · Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao ·

    ChronoVision:通过潜在状态重建实现时间推理

    arXiv:2608.05631v1 Announce Type: new Abstract: Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reas…

  197. arXiv stat.ML TIER_1 English(EN) · Cai Zhou, Zekai Wang, Menghua Wu, Qianyu Julie Zhu, Flora C. Shi, Chenyu Wang, Ashia Wilson, Tommi Jaakkola, Stephen Bates ·

    在线推理校准:测试时训练实现可泛化的共形大语言模型推理

    arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-tra…

  198. arXiv stat.ML TIER_1 English(EN) · Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho ·

    当多种答案有效时,投票失效:LLM中K选最佳因果推理的符号验证

    arXiv:2608.03506v1 Announce Type: cross Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple vali…

  199. arXiv cs.CV TIER_1 English(EN) · Jiaxuan Kang, Siyu Chen, Mingda Li, Mingjie Liu, Tianyue Wang, Zhaoyang Wei, Yongheng Zhang, Yanchao Hao, Zheng Wei ·

    LUT:视觉推理的潜在效用训练

    arXiv:2608.00743v1 Announce Type: new Abstract: Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden-space computation before answering, but they often …

  200. arXiv cs.CV TIER_1 English(EN) · Shoutai Zhu, Tianyang Xu, Bin Sun, Mingyuan Xu, Yu Liu, Qinzhen Guo ·

    OPLD:用于多模态推理的在线策略潜在蒸馏

    arXiv:2607.28154v1 Announce Type: new Abstract: Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces a…

  201. arXiv stat.ML TIER_1 English(EN) · Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Weijie J. Su, Edgar Dobriban ·

    Statistical Early Stopping for Reasoning Models

    arXiv:2602.13935v2 Announce Type: replace-cross Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introd…

  202. arXiv stat.ML TIER_1 English(EN) · Zihan Dong, Zhixian Zhang, Yang Zhou, Can Jin, Ruijia Wu, Linjun Zhang ·

    评估LLM在不知道答案时的情况:通过比较信号进行数学推理的统计评估

    arXiv:2602.03061v2 Announce Type: replace-cross Abstract: Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms. On difficult problems,…

  203. LessWrong (AI tag) TIER_1 English(EN) · Chandram Dutta ·

    终止电路(推理模型如何停止思考)。

    <p><span>Reasoning models since the dawn of o1 and R1 have a tendency to overthink. Despite a lot of work on early-exit methods and steering, open-weight and smaller reasoning models still produce long chains of thought before they answer. I worked on discovering how much of that…

  204. AWS Machine Learning Blog TIER_1 English(EN) · Rushil Anirudh ·

    探索使用 Amazon Nova 进行监督微调的自蒸馏推理

    In this post, we explore an idea for generating thinking tokens for datasets that lack reasoning traces in SFT customization. We first examine the reasoning suppression problem, then introduce Self-Distilled Reasoning (SDR), validate it across three benchmarks, and provide practi…

  205. Sequoia Capital TIER_1 English(EN) · amoore ·

    与 Etched 合作:构建推理引擎

    <p>The post <a href="https://sequoiacap.com/article/partnering-with-etched-building-the-inference-machine/">Partnering with Etched: Building the Inference Machine</a> appeared first on <a href="https://sequoiacap.com">Sequoia Capital</a>.</p>

  206. Medium — fine-tuning tag TIER_1 English(EN) · Stancho Stanchev ·

    微调 Qwen3–4B 与 SmolLM3–3B 在相同的数学推理配方上

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@stanchoz3/fine-tuning-qwen3-4b-vs-smollm3-3b-on-the-same-math-reasoning-recipe-9fad07957de3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2600/0*m0hETYuphBsrKtcb…

  207. Medium — Claude tag TIER_1 English(EN) · Nadeem Khan(NK) ·

    思维链与草稿板推理:模型大声思考

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://nadeem4-nk13.medium.com/chain-of-thought-and-scratchpad-reasoning-the-model-thinking-out-loud-d998a5a1860d?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/0*4E12i9OTIivI4pFp" …

  208. Medium — fine-tuning tag TIER_1 Bahasa(ID) · Angga Yulian Adi Pradana ·

    GRPO微调LLM以训练推理模型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anggapradanaa/grpo-fine-tuning-llm-untuk-melatih-reasoning-model-6adbf065c48a?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1734/1*xyCJt-bhgqHvLf3rS_197g.png" wi…

  209. dev.to — LLM tag TIER_1 English(EN) · Hunter G ·

    The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    <h1> The Inference Engineering Masterclass — Philip Kiely &amp; Ali Taha, Baseten </h1> <p>The same open-weights model, served by different providers, can differ by 4x to 10x in speed. So "which model" only answers half the question. The other half is whose inference.</p> <p>Base…

  210. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    ReflectRL:将失败的LLM推理转化为训练信号

    <p><em>ReflectRL introduces a novel framework that utilizes 'Golden Negative Trajectories'—failed reasoning attempts by expert models—to improve LLM performance. By treating these failures as opportunities for reflection rather than discarding them, the method enhances reasoning …

  211. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    LLM 中的推理错误并非随机噪声;它们锚定在残差流中的特定方向。识别这些高维向量

    Reasoning errors in LLMs aren't just random noise; they are anchored to specific directions within the residual stream. Identifying these high-dimensional vectors allows for targeted interventions, moving us from black-box prompting to surgical model steering. # LLMs # AI

  212. r/LocalLLaMA TIER_1 English(EN) · /u/jwdeaver ·

    LabyrinthBench:一个专注于本地、无裁判的LLM基准测试,用于衡量多步代理任务中的干扰下的上下文召回率。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vhz2rc/labyrinthbench_a_localfocused_judgefree_llm/"> <img alt="LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks." src="ht…

  213. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    从提示到推理:ToolArtist 如何统一多步逻辑与图像生成

    <h1> From Prompting to Reasoning: How ToolArtist Unifies Multi-Step Logic and Image Generation </h1> <p>Text-to-image (T2I) systems have reached a level of visual fidelity that was difficult to imagine only a few years ago. However, even the most advanced diffusion models and aut…

  214. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Claude Opus 5: ARC-AGI-3 的飞跃实际上告诉了我们关于推理进展的什么

    <h1> Claude Opus 5: What the ARC-AGI-3 Leap Actually Tells Us About Reasoning Progress </h1> <p>Anthropic released <a href="https://www.anthropic.com/news/claude-opus-5" rel="noopener noreferrer">Claude Opus 5</a> on July 24, 2026, and the headline number is hard to ignore: a 30.…

  215. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    Chain-of-Draft:保留推理过程,去除叙述,并削减约80%的推理代币

    <p>Chain-of-Thought reliably lifts reasoning accuracy by making a model write its intermediate steps down instead of leaping to an answer. But look at a CoT trace closely and you'll notice something: on a simple arithmetic problem the model emits a paragraph where the actual work…

  216. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    对比式思维链:将每个正确推理与一个标记的错误推理配对,教会模型推理边界

    <p>Plain chain-of-thought hands the model a few <em>correct</em> worked examples and hopes it generalizes the right steps. It's genuinely strong — but I kept watching it walk into the same trap on multi-step arithmetic: adding when it should subtract, jumbling the step order, gra…

  217. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    SIGMA 将确定性 AI 推理转化为透明、可追溯的语言,同时不影响可复现性或治理。了解 NEUROVATIC 的 sover

    SIGMA transforms deterministic AI reasoning into transparent, traceable language without compromising reproducibility or governance. Discover NEUROVATIC's sovereign language intelligence architecture: https:// neurovatic.ai/sigma # AI # SovereignAI # SIGMA

  218. dev.to — LLM tag TIER_1 English(EN) · Daniel Dong ·

    Kimi K3:没有“思考”开关的推理模型

    <p>A reasoning model that doesn't make you turn on "think mode."</p> <p>Kimi K3 thinks before every answer — no toggle, no config, no prompt engineering.</p> <p>1M context window. Drop a codebase. Get a review. No chunks required.<br /> </p> <div class="highlight js-code-highligh…

  219. r/LocalLLaMA TIER_1 English(EN) · /u/hellajacked ·

    MindControl - llama.cpp 分支,通过采样期间的注入来指导推理过程

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v3ms3c/mindcontrol_llamacpp_fork_to_guide_the_reasoning/"> <img alt="MindControl - llama.cpp fork to guide the reasoning process via injection during sampling" src="https://preview.redd.it/yo4avmy3dteh1.png?w…

  220. dev.to — LLM tag TIER_1 English(EN) · Satvik Mishra ·

    KNOT — 一个进行自我辩论的实验性推理架构(首个原型展示)

    <p>Meet the <em>knot</em>. I spent a month making him argue with himself.<br /> <a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.co…

  221. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    已验证的代理推理是从概率生成到可验证输出的必要转折点。通过将LLM推理锚定在工具证明的核心证据上

    Verified Agentic Reasoning is the necessary pivot from probabilistic generation to verifiable output. By anchoring LLM inferences in tool-attested kernel proofs, we move beyond "plausible" text to functional accuracy, essential for high-stakes empirical science. # AI # LLMs (1/2)

  222. dev.to — LLM tag TIER_1 English(EN) · Tech Signal Daily ·

    逻辑不一致的提示如何将推理模型变成拒绝服务问题

    <p>Reasoning models are supposed to be better at hard tasks because they “think” before answering. In practice, that extra step-by-step process creates a new attack surface: if you can push the model into overthinking, you can make it spend far more tokens than normal and slow it…

  223. r/LocalLLaMA TIER_1 English(EN) · /u/marcodsn ·

    [研究/模型] Flint:在不破坏推理能力的情况下进行压缩

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uv9o2u/studymodels_flint_compressing_reasoning_without/"> <img alt="[Study/Models] Flint: Compressing Reasoning Without Breaking It" src="https://external-preview.redd.it/oW5M6NGWwIm7UmrIEq6LxCtz9PZYMnGJKCXJx…

  224. dev.to — LLM tag TIER_1 English(EN) · jackma ·

    围绕LLM推理构建学习应用

    <p><strong>Building a Study App Around LLM Reasoning</strong></p> <p>When building an AI study app, it is easy to focus on the output: the answer, the final number, the completed explanation.</p> <p>But the more interesting product question is what happens before and around that …

  225. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Thinking Machines 发布 Inkling Small:参数量不到前代三分之一的开源推理模型,基准测试结果更佳

    Thinking Machines veröffentlicht Inkling Small: Open-Weights-Reasoning-Modell mit unter einem Drittel der Parameter des Vorgängers, aber besseren Benchmark-Ergebnissen. Effizienzgewinn durch Architektur statt reiner Skalierung. https:// the-decoder.de/thinking-machin es-setzt-mit…

  226. Mastodon — mastodon.social TIER_1 English(EN) · strike007 ·

    转向工具可验证证明暴露了当前黑盒模型的脆弱性。在我们弥合随机推理与形式验证之间的差距之前

    The shift toward tool-attested proofs exposes the fragility of current black-box models. Until we bridge the gap between stochastic reasoning and formal verification, every agentic workflow remains a liability in environments requiring audit-ready certainty. # AI # Safety (2/2)

  227. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    神经符号AI论文将逻辑规则与能量模型融合 arXiv新预印本融合答案集编程与基于能量的模型以实现端到端AI推理

    Neurosymbolic AI paper fuses logic rules with energy models A new arXiv preprint fuses Answer Set Programming with energy-based models for end-to-end AI reasoning, tested on visual QA and multi-object tracking benchmarks https://www. notatechguy.com/neurosymbolic- ai-paper-fuses-…

  228. r/singularity TIER_2 English(EN) · /u/yogthos ·

    Ring-Zero:将零样本强化学习扩展到万亿参数以实现涌现式推理

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1uyp9xe/ringzero_scaling_zero_rl_to_a_trillion_parameters/"> <img alt="Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning" src="https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnC…