PulseAugur
中
实时 07:04:24
English(EN) REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

新研究应对大语言模型推理、效率和蒸馏挑战 · 跟踪 10 个来源

新研究探索了提高大语言模型(LLMs)推理能力和效率的方法。一篇论文引入了“Trace as State”,通过将推理痕迹置于上下文块之前来增强长上下文推理,在各种模型和数据集上显示出显著的性能提升。另一项研究提出了“SALA”,用于上下文学习中的语义感知逻辑对齐,改进了复杂推理任务的演示选择。此外,关于“HEAL”的研究旨在通过解决当前蒸馏技术的局限性,从 LLMs 中将推理蒸馏到更小的模型中,其灵感来源于教育理论。 AI

影响 这些研究工作旨在提高 LLM 推理的准确性、效率和对齐能力,有望带来更强大、更可靠的 AI 系统。

排序理由 该集群包含多篇发表在 arXiv 上的学术论文,详细介绍了改进 LLM 推理和评估的新方法和分析。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 296 个来源。 我们如何撰写摘要 →

新研究应对大语言模型推理、效率和蒸馏挑战 · 跟踪 10 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含多篇发表在 arXiv 上的学术论文,详细介绍了改进 LLM 推理和评估的新方法和分析。
Source corroboration
296 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
60 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+94 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [296]

  1. arXiv cs.CL TIER_1 English(EN) · Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong ·

    反思性恢复:一种通过从错误中学习进行推理的自监督方法

    arXiv:2609.19156v1 Announce Type: new Abstract: Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajector…

  2. arXiv cs.AI TIER_1 English(EN) · Chengwen Qi, Deheng Ye, Yatao Bian ·

    当前的系统性泛化任务错过了什么?一项以推理为中心的分析

    arXiv:2609.19212v1 Announce Type: new Abstract: Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on si…

  3. arXiv cs.AI TIER_1 English(EN) · Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak ·

    When2Think:学习面向效率的混合推理模型中的难度感知长度控制

    arXiv:2609.19671v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rig…

  4. arXiv cs.AI TIER_1 English(EN) · Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar, Zachary Ives ·

    高效链接非结构化数据以实现多步推理

    arXiv:2609.19491v1 Announce Type: cross Abstract: Modern LLMs and AI agents increasingly support data engineering workflows that integrate evidence from unstructured sources. Such pipelines typically do data retrieval, integration, and ranking before proceeding to more complex ag…

  5. arXiv cs.AI TIER_1 English(EN) · Srijith Ravikumar ·

    缺失的“我不知道”:为何三个推理可靠性发现汇聚于校准性弃权

    arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymanov et al. (2026) show that under safety-constrained generatio…

  6. arXiv cs.AI TIER_1 English(EN) · Anish Sathyanarayanan, Aditya Nagarsekar, Aarush Rathore ·

    绕过合理性:语言模型隐式推理的因果审计

    arXiv:2602.03994v3 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) prompting is widely used as a reasoning aid and is often treated as a transparency mechanism. Yet behavioral gains under CoT do not imply that the model's internal computation causally depends on the…

  7. arXiv cs.CL TIER_1 English(EN) · Magnus Boman ·

    理解大型语言模型(LLM)的失败:基于多磁带图灵机的系统性推理错误分析

    arXiv:2602.15868v3 Announce Type: replace Abstract: Large language models (LLMs) exhibit failure modes on seemingly trivial tasks. We propose a formalisation of LLM interaction using a deterministic multi-tape Turing machine, where each tape represents a distinct component: input…

  8. arXiv cs.AI TIER_1 English(EN) · Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran, Dung D. Le ·

    OBC-Prune:面向大型推理模型剪枝的基于结果的校准

    arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate para…

  9. arXiv cs.AI TIER_1 English(EN) · Cai Ke, Xinghao Chen, Xiaoyu Shen, Keyu Chen, Siyu An, Junnan Dong, Ruifeng Xu, Ruizhi Qiao, Xing Sun ·

    通过潜在神经符号推理解开长期记忆

    arXiv:2609.18461v1 Announce Type: new Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence. While early flat retrieval methods score memory fragments independently and neglect …

  10. arXiv cs.AI TIER_1 English(EN) · Yizheng Yang, Haining Yu, Yuechen Wang, Yikai Hou, Xing Fu, Jinbo Yang, Tianqing Zhu ·

    首个 Token 至关重要:理解大型推理模型中的安全崩溃

    arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preferen…

  11. arXiv cs.AI TIER_1 English(EN) · Rebecca Ansell, Autumn Toney-Wails ·

    利用工具增强的演绎推理来启发大型语言模型

    arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that require integrating evidence across multiple reasoning steps, ma…

  12. arXiv cs.AI TIER_1 English(EN) · Zhongdi Qu, Carla P. Gomes ·

    大型语言模型数学推理中的四阶段分解和机械脆弱性问题

    arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these observations with a mechanistic account. We show that the m…

  13. Hugging Face Daily Papers TIER_1 English(EN) ·

    When2Think:学习难度感知长度控制,用于高效混合推理模型

    Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading redu…

  14. arXiv cs.AI TIER_1 English(EN) · Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan ·

    大型语言模型能否模拟学生错误推理?以干扰项生成为例

    arXiv:2603.15547v2 Announce Type: replace-cross Abstract: Modeling student misconceptions in a realistic manner is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating distractor answers for multiple-c…

  15. arXiv cs.AI TIER_1 English(EN) · Zhiren Gong, Yikun Hou, Zihao Zeng, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim ·

    State of Thought 促进内源性推理

    arXiv:2609.16055v1 Announce Type: cross Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reas…

  16. arXiv cs.AI TIER_1 English(EN) · Hung-Hsuan Chen ·

    深度递归Transformer实现高效测试时推理,以深度而非长度实现组合泛化

    arXiv:2603.21676v2 Announce Type: replace-cross Abstract: Standard Transformers have a fixed computational depth, limiting their ability to generalize to tasks that require variable-depth reasoning. The usual remedy, Chain-of-Thought (CoT), spends tokens to reason, inflating the …

  17. arXiv cs.AI TIER_1 English(EN) · Jinyang Zhang, Weibin Liao, Keqin Bao, Sihang Li, Shaobo Wang, Muyang Ye, Hongxin Ding, Yue Fang, Tianyi Tang, Fei Huang, Kexin Yang, Xingzhang Ren, Dayiheng Liu ·

    模仿游戏:当大型语言模型通过以代码为中心的推理数据合成学会像程序一样推理

    arXiv:2609.16076v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge …

  18. arXiv cs.AI TIER_1 English(EN) · Pratibha Zunjare, Michael Hsiao ·

    NeuroProlog:通过鸡尾酒效应实现神经符号数学推理的多任务微调

    arXiv:2603.02504v3 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on natural language tasks but remain unreliable in mathematical reasoning, frequently generating fluent yet logically inconsistent solutions. We present \textbf{NeuroProlog…

  19. arXiv cs.AI TIER_1 English(EN) · Hongyu Gu, Chang Liu, Jingwen Fu ·

    向量能容纳多少想法?叠加推理的容量

    arXiv:2609.13747v1 Announce Type: new Abstract: Large language models solve hard problems through intermediate computations across multi-step reasoning. Traditional chain-of-thought encodes these computations as tokens. Recent continuous and recurrent methods instead move partial…

  20. arXiv cs.AI TIER_1 English(EN) · Simon Schug, Brenden M. Lake ·

    无系统性思考?在规则归纳任务上评估推理模型

    arXiv:2609.13948v1 Announce Type: cross Abstract: A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so…

  21. arXiv cs.CL TIER_1 English(EN) · Fangan Dong, Zuming Yan, Xuri Ge, Zhiwei Xu, Mengqi Zhang, Xuanang Chen, Ben He, Xin Xin, Zhumin Chen, Ying Zhou ·

    识别和转移关键推理神经元:通过激活引导提高 LLM 推理的可靠性

    arXiv:2601.19847v3 Announce Type: replace Abstract: Despite the strong reasoning capabilities of recent large language models (LLMs), achieving reliable performance on challenging tasks often requires post-training or computationally expensive sampling strategies, limiting their …

  22. arXiv cs.CL TIER_1 English(EN) · Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar, Yannick Detrois, Mattia Nee, Pierre-Carl Langlais, Ivan P. Yamshchikov ·

    谁来评估评估基准?迈向常识推理基准的全面评估

    arXiv:2504.07825v2 Announce Type: replace Abstract: Commonsense reasoning is a key language model capability, as it is purportedly a prerequisite for many basic tasks, unlike specific factual knowledge. It is often measured with multiple-choice questions (MCQ) benchmarks, e.g. He…

  23. arXiv cs.CL TIER_1 English(EN) · Emmy Liu, Graham Neubig, Jacob Andreas ·

    未竟的循环:语言模型中的演绎、归纳和溯因推理

    arXiv:2404.03028v4 Announce Type: replace Abstract: Modern language models (LMs) can learn to perform new tasks in different ways: in instruction following, the target task is described explicitly in natural language; in few-shot prompting, the task is specified implicitly with a…

  24. arXiv cs.CL TIER_1 English(EN) · Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada, Sara Khosravi, Zeynep Senahan Yildiz, Amit Sahai, Nanyun Peng ·

    推理能否提升大型语言模型的心理深度?这取决于评判者

    arXiv:2609.13773v1 Announce Type: cross Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human pre…

  25. arXiv cs.CL TIER_1 English(EN) · Runa Yoshida, Kosuke Nishida, Kyosuke Nishida ·

    通过推理过程错误分类提升大型语言模型数学推理能力

    arXiv:2609.15145v1 Announce Type: new Abstract: The reasoning ability of large language models (LLMs) is a critical factor for practical LLM-based applications. To investigate the current reasoning capability of LLMs, we clarify the types of errors that arise in LLMs' reasoning p…

  26. arXiv cs.AI TIER_1 English(EN) · Danchun Chen, Qiyao Yan, Chenpeng Wang, Liangming Pan ·

    大型语言模型中命题逻辑推理的机械化理解探究

    arXiv:2601.04260v2 Announce Type: replace Abstract: Understanding how Large Language Models (LLMs) perform logical reasoning internally remains a fundamental challenge. While prior mechanistic studies focus on identifying task specific circuits, they leave open the question of wh…

  27. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Adam Jatowt ·

    幅度幻觉:重新思考推理密集型检索的置信度

    Many production RAG systems implement retrieval abstention by thresholding raw similarity scores, implicitly treating score magnitude as a confidence signal. We demonstrate that this practice degrades systematically as queries require reasoning beyond semantic matching. Across 11…

  28. arXiv cs.AI TIER_1 English(EN) · Yasmine Briefs, Christoph Weidenbach ·

    通过非地面子句学习扩展SMT求解器

    arXiv:2609.11509v1 Announce Type: new Abstract: Quantifier instantiation is currently the main approach to non-ground SMT solving: solvers generate ground instances and solve the resulting ground SMT problems with CDCL(T)-style reasoning. When a conflict is found, conflict analys…

  29. Hugging Face Daily Papers TIER_1 English(EN) ·

    无系统性思考?在规则归纳任务上评估推理模型

    Reasoning models frequently fail on structurally equivalent task variants, indicating a lack of systematicity in their cognitive abilities.

  30. arXiv cs.AI TIER_1 English(EN) · Yiqi Li, Xu Chen, Chen Ju, Jiangchao Yao, Zhaoyang Li, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Yu Wang ·

    面向潜在思维链推理的结构化过程监督

    arXiv:2609.09928v1 Announce Type: new Abstract: Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack direct process supervisi…

  31. arXiv cs.LG TIER_1 English(EN) · Deblina Kar ·

    一种用于组合式和可解释认知推理的多阶段规则链式框架

    arXiv:2609.10654v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a multi-stage rule-chaining framework that performs compositi…

  32. arXiv cs.CL TIER_1 English(EN) · Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng ·

    负面自蒸馏:通过避免缺陷来学习推理

    arXiv:2609.11699v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. …

  33. arXiv cs.AI TIER_1 English(EN) · Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao, Xiaoyuan Wang, Soroush Vosoughi ·

    SFT 实际应该学习哪些 Token?从 Token 修剪视角看数学推理

    arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already master…

  34. Hugging Face Daily Papers TIER_1 English(EN) ·

    负面自蒸馏:通过避免缺陷来学习推理

    On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can …

  35. arXiv cs.AI TIER_1 English(EN) · Kang Chen, Sihan Zhao, Yixin Cao, Yu-Gang Jiang ·

    从集中到差异化再回归:MoE推理群组中的有效排名路由

    arXiv:2609.06403v1 Announce Type: new Abstract: Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effectiv…

  36. arXiv cs.AI TIER_1 English(EN) · Qihao Yuan ·

    Deposon:一个可审计、保证守恒、经过博弈论测试的 LLM 推理路径散射层

    arXiv:2609.09001v1 Announce Type: new Abstract: Multi-step LLM reasoning lacks a machine-recheckable ledger: discarded reasoning paths leave no auditable record. We propose the Deposon scattering layer, which binds each node of an LLM-generated concept-decomposition graph to a tw…

  37. arXiv cs.AI TIER_1 English(EN) · Yu-Hang Wu, Yu-Jie Xiong, Henghua Zhang, Bairui Zhang, Jia-Chen Zhang, Shaohua Li ·

    更深层次的推理会损害对齐吗?大型推理模型中对齐崩溃的揭示与缓解

    arXiv:2609.08186v1 Announce Type: new Abstract: The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed to enhance safety alignment, the stability of alignment mechanisms under extended…

  38. arXiv cs.AI TIER_1 English(EN) · Suhyeong Park, Junha Jung, Jaewoo Kang ·

    推理潜藏之中!让潜藏视觉推理成为必要

    arXiv:2609.06746v1 Announce Type: new Abstract: Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model …

  39. arXiv cs.LG TIER_1 English(EN) · Rafael Pardinas, Ehsan Kamalloo, David Vazquez, Alexandre Drouin ·

    Apriel-Reasoner:用于通用高效推理的 RL 后训练

    arXiv:2604.02007v3 Announce Type: replace Abstract: Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopted by frontier open-weight models. However, their training recipes and domain mixt…

  40. arXiv cs.LG TIER_1 Dansk(DA) · Baris Arat, Emre Sefer ·

    固定证据池下LLM重排器行为诊断

    arXiv:2602.18613v2 Announce Type: replace Abstract: Standard reranking evaluations study how a reranker orders candidates returned by an upstream retriever. This setup couples ranking behavior with retrieval quality, so differences in output cannot be attributed to the ranking po…

  41. arXiv cs.LG TIER_1 English(EN) · Jaewoo Lim, Sungbok Shin, Sanghyun Hong ·

    推理表征能否帮助人类评估LLM输出?

    arXiv:2609.09038v1 Announce Type: new Abstract: Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether the…

  42. arXiv cs.LG TIER_1 English(EN) · Yuwen Hao, Menglin Yang ·

    拓宽思路:缓解隐式思维链推理中的潜在秩崩溃

    arXiv:2609.07406v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rationales increase decoding length, latency, and context cost. Implicit CoT offers a mor…

  43. arXiv cs.LG TIER_1 English(EN) · Yimeng Ye, Shuang Chen, Wenxuan Huang, Manyuan Zhang, Kaituo Feng, Zhangquan Chen, Jiayu Chen, Yucheng Zhou, Yicheng Xiao, Zhiyuan Feng, Tianyu Shi ·

    Stable-MM-R1:通过熵引导分层锚定多模态推理动态

    arXiv:2609.07148v1 Announce Type: new Abstract: While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" …

  44. arXiv cs.CL TIER_1 English(EN) · Ryan Lail ·

    分解LLM-裁判不确定性以定位专家标签

    arXiv:2609.06444v2 Announce Type: replace Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, a…

  45. arXiv cs.CL TIER_1 English(EN) · Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu ·

    ConvMem:用于长上下文推理的卷积记忆

    arXiv:2609.10441v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective …

  46. arXiv cs.AI TIER_1 English(EN) · Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang ·

    通过受人类启发的奖励塑造来增强 LLM 推理能力

    arXiv:2602.04265v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, existing reward formulations typically treat exploration and conso…

  47. arXiv cs.AI TIER_1 English(EN) · Zuoyou Jiang, Li Zhao, Rui Sun, Ruohan Sun, Zhongjian Li, Jing Li, Daxin Jiang, Zuo Bai, Cheng Hua ·

    Alpha-R1:通过强化学习利用LLM推理进行Alpha筛选

    arXiv:2512.23515v2 Announce Type: replace-cross Abstract: Signal decay and regime shifts pose recurring challenges for data-driven investment strategies in non-stationary markets, where conventional time-series and machine learning approaches often struggle to generalize beyond h…

  48. arXiv cs.AI TIER_1 Deutsch(DE) · Yufeng Zhao, Junnan Liu, Hongwei Liu, Dongsheng Zhu, Yuan Shen, Songyang Zhang, Kai Chen ·

    当工具损害大语言模型推理能力:外部证据下的状态依赖信念修正

    arXiv:2508.15754v2 Announce Type: replace-cross Abstract: Tool use is often assumed to monotonically improve reasoning, where external evidence is expected to help when relevant and be ignored when irrelevant. We show that this assumption fails in a state-dependent way. Across be…

  49. arXiv cs.AI TIER_1 English(EN) · Andreas Reich, Claudia Thoms, Tobias Schrimpf ·

    推出 HALC:用于计算社会科学中 LLM 自动编码的系统可靠提示构建通用流水线

    arXiv:2507.21831v2 Announce Type: replace-cross Abstract: LLMs are seeing widespread use for task automation, including automated coding in the social sciences. However, even though researchers have proposed different prompting strategies, their effectiveness varies across LLMs a…

  50. arXiv cs.AI TIER_1 English(EN) · Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder ·

    当智能体言行不一时:验证从大型语言模型中提取的信念

    arXiv:2602.06286v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of unknown outcomes. However, it is unclear whether LLMs act as if they hold cohere…

  51. arXiv cs.AI TIER_1 English(EN) · Sining Zhoubian, Dan Zhang, Jie Tang ·

    ReST-RL:通过统一的自训练和价值引导搜索来增强LLM推理能力

    arXiv:2508.19576v3 Announce Type: replace Abstract: With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method - Group Relative Policy Optimization (GRPO) - has achieved critical success, yet it still suffers from the issue of …

  52. arXiv cs.AI TIER_1 English(EN) · Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He ·

    A*-Thought-V2:通过LLM的几何动力学实现高效的潜在推理

    arXiv:2609.07821v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a princ…

  53. arXiv cs.AI TIER_1 English(EN) · Xiaodong Wang, Peixi Peng ·

    Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning

    arXiv:2609.07036v1 Announce Type: cross Abstract: We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations such as I'm doing, in contrast to the revision- and backtracking-oriented Aha Moment. We refer to their corresponding l…

  54. arXiv cs.AI TIER_1 English(EN) · Qi Wang, Chengcheng Wan, Jiangtao Wang ·

    SRD-GUARD:通过语义重写和联合多模型评分暴露潜在意图的LLM防御框架

    arXiv:2609.06540v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inferenc…

  55. arXiv cs.AI TIER_1 English(EN) · Mar Gonz\`alez I Catal\`a, Haitz S\'aez de Oc\'ariz Borde, Davide Murari, Carola-Bibiane Sch\"onlieb, Pietro Li\`o, George Monta\~nez ·

    Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

    arXiv:2609.09030v1 Announce Type: new Abstract: Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work …

  56. Hugging Face Daily Papers TIER_1 English(EN) ·

    负面自蒸馏:通过避免缺陷来学习推理

    Negative Self-Distillation improves large language model reasoning by pushing models away from self-generated flawed reasoning via a dynamic gating mechanism that protects linguistic capabilities.

  57. Hugging Face Daily Papers TIER_1 English(EN) ·

    ConvMem:用于长上下文推理的卷积记忆

    While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iterativel…

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    A*-Thought-V2:通过LLM的几何动力学实现高效的潜在推理

    A*-Thought-V2 models chain-of-thought reasoning as hidden-state trajectories to selectively retain explicit reasoning steps or compress them into continuous latent tokens, improving accuracy and efficiency.

  59. arXiv cs.CL TIER_1 English(EN) · Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu, Alice Oh, Taekyung Kim ·

    深入理解思维链:大型语言模型推理操作的机制化解读

    arXiv:2609.04753v1 Announce Type: new Abstract: Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about …

  60. arXiv cs.AI TIER_1 English(EN) · Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri ·

    必要还是充分?用行为证据评估 LLM 的解释

    arXiv:2609.05385v1 Announce Type: new Abstract: LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or dec…

  61. arXiv cs.AI TIER_1 English(EN) · Andrea Gregor de Varda, Sana Pandey, Pengrui Han, Jacob Andreas, Evelina Fedorenko ·

    共享电路预测大型语言模型是否能跨格式泛化算术推理

    arXiv:2609.04463v1 Announce Type: cross Abstract: In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solve 'two plus five'. In contrast, LLMs are more brit…

  62. arXiv cs.AI TIER_1 English(EN) · Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane ·

    极度稀疏的监督激励推理能力

    arXiv:2609.04565v1 Announce Type: new Abstract: Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning…

  63. arXiv cs.AI TIER_1 English(EN) · Weicai Huang (Beijing MQPat Technologies, Co., Ltd.) ·

    DODR: 潜在空间中的确定性算子驱动推理

    arXiv:2609.04782v1 Announce Type: new Abstract: Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which induces three fundamental defects in complex logical reasoning: error accumulation, probability substituting necessity, and t…

  64. arXiv cs.AI TIER_1 English(EN) · Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou, Shao-Qun Zhang ·

    GUT:通过图复杂度量化和优化LLM的推理不确定性

    arXiv:2609.05284v1 Announce Type: new Abstract: Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs). However, the reasoning processes of LLMs often exhibit uncertainty, where LLMs often produce a proliferation of divergent branches …

  65. arXiv cs.LG TIER_1 English(EN) · Jeffrey Lai, Anthony Bao, John Quinn, William Gilpin ·

    分形盆地捕获潜在推理

    arXiv:2609.04963v1 Announce Type: new Abstract: Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widel…

  66. arXiv cs.CL TIER_1 English(EN) · Shi-Qi Yan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling ·

    ConsensusBench: 通过结果奖励致密化进行LLM推理的共识节点基准测试

    arXiv:2609.04648v1 Announce Type: new Abstract: Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong …

  67. Hugging Face Daily Papers TIER_1 English(EN) ·

    重新审视训练后完整推理轨迹

    Large language models gain reasoning improvements from truncated trajectory endpoints rather than full reasoning traces, reducing redundancy while benefiting supervised fine-tuning and reinforcement learning.

  68. Hugging Face Daily Papers TIER_1 English(EN) ·

    洞悉潜在!使潜在视觉推理成为必要

    CVRR enforces recurrent hidden-state computation as the required image-conditioned pathway for visual reasoning, preserving model competence while distinguishing latent information from actual predictive use.

  69. arXiv cs.AI TIER_1 English(EN) · Yigit Utku Bulut ·

    问题所在,而非路径:LLM推理轨迹中的预算与难度困扰

    arXiv:2609.03436v1 Announce Type: cross Abstract: Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both …

  70. arXiv cs.LG TIER_1 English(EN) · Yuxiang Wang, Kunyu Feng, Yingda Shen, Haoning Xu, Junyu Wang, Zhizheng Wu ·

    RecurTrace:具有循环时间记忆的自适应潜在推理

    arXiv:2609.03379v1 Announce Type: new Abstract: Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two…

  71. arXiv cs.LG TIER_1 English(EN) · Leqi Zheng, Jinbo Su, Fang Niu, Chaokun Wang, Weiping Wang, Jiajun Zhang, Shannan Yan, Jie Wu, Zhaolu Kang, Rong Fu, Hang Zhang ·

    梯度知晓结果不知:利用梯度对齐奖励解锁LLM推理的强化学习

    arXiv:2609.03342v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from …

  72. arXiv cs.LG TIER_1 English(EN) · Frank Hu, Shriram Chennakesavalu, David Graff ·

    Frontier LLMs 是有效的批量优化器:在连续和离散环境中评估推理模型

    arXiv:2609.03177v1 Announce Type: new Abstract: Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoni…

  73. arXiv cs.CL TIER_1 English(EN) · Harshad Khadilkar, Abhay Gupta ·

    Causal-Counterfactual RAG:将因果反事实推理整合到RAG中

    arXiv:2509.14435v3 Announce Type: replace Abstract: Large language models (LLMs) have transformed natural language processing (NLP), enabling diverse applications by integrating large-scale pre-trained knowledge. However, their static knowledge limits dynamic reasoning over exter…

  74. arXiv cs.AI TIER_1 English(EN) · Hui Wu, Hengyi Cai, Jinman Zhao, Xinran Chen, Ziheng Li, Zhejun Zhao, Shuaiqiang Wang, Yuchen Li, Dawei Yin ·

    并非所有偏好都值得梯度:理解离线推理对齐中的梯度效用

    arXiv:2602.01207v2 Announce Type: replace Abstract: Offline preference optimization aligns reasoning models from fixed chosen--rejected pairs, yet standard methods apply gradient updates from every pair regardless of its training value under the current policy. We argue that this…

  75. arXiv cs.AI TIER_1 English(EN) · Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Chu Liu, Pengjun Xie, Yilun Zhao, Shu Wu ·

    CORE:通过重排序器蒸馏改进多模态大模型嵌入中的组合推理

    arXiv:2609.04083v1 Announce Type: cross Abstract: MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions w…

  76. arXiv cs.AI TIER_1 English(EN) · Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim ·

    </think> 仍能进行推理:对伪造的 CoT 终止的分析

    arXiv:2609.03633v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to …

  77. Hugging Face Daily Papers TIER_1 English(EN) ·

    深入理解思维链:大型语言模型推理操作的机制化解读

    Distinct reasoning operations in language models are geometrically separable in hidden representations, with structure emerging across layers and depending on contextual reasoning context.

  78. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Shu Wu ·

    CORE:通过重排序器蒸馏改进多模态大模型嵌入中的组合推理

    MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating…

  79. Hugging Face Daily Papers TIER_1 English(EN) ·

    </think> 仍能推理:对伪造 CoT 终止的分析

    Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that in…

  80. Hugging Face Daily Papers TIER_1 English(EN) ·

    梯度知晓结果不知:通过梯度对齐奖励解锁LLM推理的强化学习

    Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, eit…

  81. arXiv cs.CL TIER_1 English(EN) · Xu Zou, Jie Tang ·

    Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

    arXiv:2609.02702v1 Announce Type: new Abstract: Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing t…

  82. arXiv cs.AI TIER_1 English(EN) · Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld, Joshua Tenenbaum, Kevin Ellis ·

    通过语言和代码的概率推理进行归纳和探究

    arXiv:2609.01815v1 Announce Type: new Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) d…

  83. arXiv cs.AI TIER_1 English(EN) · Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal, Robert Kirk, John Hughes ·

    通过推理时计算和部署脚手架提高评估的真实性

    arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make…

  84. arXiv cs.AI TIER_1 English(EN) · Zhao Ji, Wenqing Chen, Zhixuan Chu, Jianxing Yu, Jingping Liu, Shanhe Zhao, Zibin Zheng ·

    SALA:在上下文学习中实现复杂推理的语义感知逻辑对齐

    arXiv:2609.02336v1 Announce Type: new Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-bas…

  85. arXiv cs.AI TIER_1 English(EN) · Rongzhi Zhu, Yi Liu, Jiancheng Wang, Xiangyu Liu, Zequn Sun, Yiwei Wang, Yu Deng, Zijian Zhou, Wei Hu ·

    大型推理模型何时能拯救思考?行为分歧的机制分析

    arXiv:2505.15276v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to "overthink" leads to inefficiencies. Although "save-thinking" prompts are intended to mitigate this issue, we find that LRMs …

  86. Hugging Face Daily Papers TIER_1 English(EN) ·

    CORE:通过重排序器蒸馏改进多模态大模型嵌入中的组合推理

    CORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.

  87. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过推理时计算和部署脚手架提高评估的真实性

    A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to disti…

  88. arXiv cs.CL TIER_1 English(EN) · Zabir Al Nazi, Shubhashis Roy Dipta, Sudipta Kar ·

    DAGGER:用于数学问题可执行推理的注意力分散感知图生成

    arXiv:2601.06853v3 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting is widely adopted for mathematical problem solving, including in low-resource languages, yet its behavior under irrelevant context remains underexplored. To systematically study this challenge, w…

  89. arXiv cs.LG TIER_1 English(EN) · Ngoc-Hieu Nguyen, Parshin Shojaee, Phuc Minh Nguyen, Nan Zhang, Chandan K Reddy, Khoa D Doan, Rui Zhang ·

    为什么推理模型会失去覆盖范围?数据和岔路口的作用

    arXiv:2605.17026v2 Announce Type: replace Abstract: Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedures. While these methods reliably improve pass@1 …

  90. arXiv cs.LG TIER_1 English(EN) · Mariia Drozdova, Aidan Sirbu, Pietro Miotti, Robert Obryk, Mayalen Etcheverry, Eyvind Niklasson, Blake Richards ·

    Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

    arXiv:2609.01449v1 Announce Type: new Abstract: Diffusion models and recursive reasoners are both iterative, but they carry information across iterations differently. We add a persistent hidden state to a diffusion denoiser and remove its timestep conditioning, leaving a single s…

  91. arXiv cs.CL TIER_1 English(EN) · Jonathan Zheng, Zirui Shao, Alan Ritter, Wei Xu ·

    用于LLM时序评估和知识更新的合成世界

    arXiv:2609.00184v1 Announce Type: new Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactu…

  92. arXiv cs.AI TIER_1 English(EN) · Jingcheng Yu, Mingliang Zeng, Qiwei Ye ·

    FloydNet:面向全局关系推理的学习范式

    arXiv:2601.19094v3 Announce Type: replace-cross Abstract: Learning algorithmic computation often requires explicit relational intermediate states, yet many graph processors maintain their primary states on individual entities. We introduce \fnet and \textbf{Pivotal Attention} (PA…

  93. arXiv cs.AI TIER_1 English(EN) · Wenjing Zhang, Jiangze Yan, Jieyun Huang, Yi Shen, Shuming Shi, Ping Chen, Ning Wang, Zhaoxiang Liu, Kai Wang, Shiguo Lian ·

    HEAL:基于事后熵辅助学习的推理蒸馏

    arXiv:2603.10359v2 Announce Type: replace Abstract: Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitations of rejection sampling. Standard methods treat the teacher as a static filter, discarding comple…

  94. arXiv cs.AI TIER_1 English(EN) · Ye Tian, Zihao Wang, Onat Gungor, Xiaoran Fan, Tajana Rosing ·

    LifeAgentBench:面向长时域、跨维度生活方式健康推理的 LLM 评测基准

    arXiv:2601.13880v2 Announce Type: replace Abstract: Personalized lifestyle health analysis requires long-horizon, multi-dimensional reasoning over heterogeneous lifestyle signals, and recent advances in mobile sensing and large language models (LLMs) make such support increasingl…

  95. arXiv cs.AI TIER_1 English(EN) · Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson ·

    HugAgent:个人级别推理的人类模拟基准

    arXiv:2510.15144v4 Announce Type: replace Abstract: Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus…

  96. arXiv cs.AI TIER_1 English(EN) · Hamed Babaei Giglou, Jennifer D'Souza, S\"oren Auer ·

    通用的自然语言处理嵌入能捕捉本体论推理吗?

    arXiv:2609.00177v1 Announce Type: cross Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings disting…

  97. arXiv cs.AI TIER_1 English(EN) · Zhaoliang Chen, Jie Fu ·

    潜在循环思想:使用冻结的LLM进行推理的提议潜在因素的循环细化

    arXiv:2609.01117v1 Announce Type: new Abstract: Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - …

  98. arXiv cs.AI TIER_1 English(EN) · Lu Cheng ·

    摆脱冗余推理:面向推理时LLM的结构感知搜索

    arXiv:2609.00738v1 Announce Type: new Abstract: Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored---a failure mode we call \textit{reasoning basin coll…

  99. arXiv cs.AI TIER_1 English(EN) · Beidi Zhao, Gexin Huang, Ciro Zhang, Anqi Li, Yusheng Tan, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li ·

    SlideBank:用于一致性全切片推理的持久化分层证据库

    arXiv:2609.00342v1 Announce Type: new Abstract: Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WS…

  100. Hugging Face Daily Papers TIER_1 English(EN) ·

    摆脱冗余推理:面向推理时LLM的结构感知搜索

    Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored---a failure mode we call \textit{reasoning basin collapse}. We introduce BASIN, a training-free, stru…

  101. arXiv cs.LG TIER_1 English(EN) · Xin Jiang, Minhao Wang, Wen Wu, Zhentao Xie, Shangheng Du, Jinxin Shi, Jiabao Zhao ·

    ERR+: 顺序熵解析,实现高效且决定性的LLM推理

    arXiv:2608.28771v1 Announce Type: new Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong resul…

  102. arXiv cs.CL TIER_1 English(EN) · Zeming Chen, Angelika Romanou, Gail Weiss, Antoine Bosselut ·

    PERK:长上下文推理作为测试时学习

    arXiv:2507.06415v3 Announce Type: replace Abstract: Long-context reasoning requires accurately identifying relevant information in extensive, noisy input contexts. In this work, we propose PERK (Parameter Efficient Reasoning over Knowledge), a scalable approach for learning to en…

  103. arXiv cs.AI TIER_1 English(EN) · Swapnil Parekh, Naman Goyal ·

    放下姿态:探测器过滤的强化学习用于忠实的思维链推理

    arXiv:2605.11467v2 Announce Type: replace-cross Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of *reasoning theater*: deliberative-looking steps that contribute nothing to correctness. This wastes inference tok…

  104. arXiv cs.AI TIER_1 English(EN) · Rubing Chen, Jian Wang, Wenjie Li, Xiao-Yong Wei, Qing Li ·

    检索还是思考?多跳复杂推理的跨边界上下文演化

    arXiv:2601.08747v3 Announce Type: replace-cross Abstract: Current context augmentation methods, such as retrieval-augmented generation, play a crucial role in bridging a model's internal knowledge boundary and external evidence for multi-hop reasoning. However, they often follow …

  105. arXiv cs.AI TIER_1 English(EN) · Yadong Wang, Haodong Chen, Yu Tian, Chuanxing Geng, Dong Liang, Xiang Chen ·

    超越稠密态:稀疏转码器作为LLM潜在推理的可因果检验算子

    arXiv:2602.01695v2 Announce Type: replace Abstract: Latent reasoning reduces the token-generation cost of chain-of-thought reasoning by replacing explicit intermediate tokens with continuous latent transitions. However, existing latent reasoning methods usually rely on dense and …

  106. arXiv cs.AI TIER_1 English(EN) · Qingjie Zhang, Yujia Fu, Yang Wang, Liu Yan, Tao Wei, Ke Xu, Minlie Huang, Han Qiu ·

    停止在失败之前:大型推理模型的运行能力边界以减轻非生产性推理

    arXiv:2509.24711v4 Announce Type: replace Abstract: Current answering paradigms for Large Reasoning Models (LRMs) often fail to account for the fact that some questions may lie beyond the model's operational capability boundary, leading to long but unproductive reasoning. In this…

  107. arXiv cs.AI TIER_1 English(EN) · Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi ·

    超越表面形式:符号编辑作为LLM逻辑推理的测试

    arXiv:2608.30256v1 Announce Type: cross Abstract: Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has oft…

  108. arXiv cs.AI TIER_1 English(EN) · Yangsong Lan, Renkai Hu, HongKai Zheng, Bo Zhang, Renzhi Wang, Hongliang Dai, Piji Li ·

    MI-Distillation:从模型插值指令推理数据谱中选择用于思维链蒸馏的方法

    arXiv:2608.29623v1 Announce Type: cross Abstract: Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains chall…

  109. arXiv cs.AI TIER_1 English(EN) · Dylan Jayabahu, Tinuade Adeleke ·

    停止向量:内部化因果引导干预以实现高效推理

    arXiv:2608.28859v1 Announce Type: cross Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable v…

  110. arXiv cs.AI TIER_1 English(EN) · Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos ·

    通过自适应结构化非结构化数据实现令牌高效数据推理代理

    arXiv:2608.31082v1 Announce Type: new Abstract: Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions fo…

  111. arXiv cs.LG TIER_1 English(EN) · Zimo Shi, Xander Tifft, Wen Xing ·

    推理模型中隐藏指令的选择性披露:行为不对称与引导

    arXiv:2608.29070v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. This assumes CoT surface…

  112. arXiv cs.AI TIER_1 English(EN) · Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo ·

    超越人类监督扩展大型推理模型:通往超级智能的道路

    arXiv:2608.31075v1 Announce Type: new Abstract: Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extendi…

  113. arXiv cs.AI TIER_1 English(EN) · Jayanta Sadhu, Sayem Shahad, Kenneth Marino ·

    DERELAB:利用生成式基准测试探究大型语言模型中的可撤销推理和确认偏误

    arXiv:2608.30413v1 Announce Type: new Abstract: Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in …

  114. arXiv cs.AI TIER_1 English(EN) · Hanlin Tian, Minhao Li, Yu Mi, Sihan Zhu, Zhao Yang, Yuxiang Wang, Hongquan Zhu, Qiufei Hu ·

    无知还是无能?为LLM代理构建知识门控、可验证的任务

    arXiv:2608.30322v1 Announce Type: new Abstract: Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that…

  115. arXiv cs.AI TIER_1 English(EN) · Yu Li, Wei Li, Xin Gao, Mengyuan Sun, Xiaoyang Wang, Qizhi Pei, Lijun Wu ·

    SPARK:基于骨架的从大规模科学文献中推理合成

    arXiv:2608.30214v1 Announce Type: new Abstract: Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited e…

  116. arXiv cs.CL TIER_1 English(EN) · Mengdan Zhu, Senhao Cheng, Liang Zhao ·

    分解、观察与推理:用于VLMs的增强潜在推理

    arXiv:2604.07518v2 Announce Type: replace Abstract: Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insuf…

  117. arXiv cs.CL TIER_1 English(EN) · Manas Mehta, Fangcong Yin, Greg Durrett ·

    Randomized YaRN 改进长上下文推理的长度泛化能力

    arXiv:2606.23687v2 Announce Type: replace Abstract: Large language models (LLMs) are typically pretrained on short sequences and then extended to work on longer sequences with additional training. However, such LLMs still struggle to further generalize to very long sequences. We …

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越人类监督扩展大型推理模型:通往超级智能的道路

    Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    DERELAB:使用生成式基准测试探究大型语言模型中的可驳回推理和确认偏见

    Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in defeasible reasoning, the datasets have been sta…

  120. arXiv cs.LG TIER_1 English(EN) · ShengYun Peng, Pin-Yu Chen, Eric Smith, Song Jiang, Hongyuan Zhan, Haozhu Wang, Mahesh Pasupuleti, Duen Horng Chau, Jianfeng Chi ·

    大型推理模型从错误思考中学习更好的对齐

    arXiv:2510.00938v3 Announce Type: replace Abstract: Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flaw…

  121. arXiv cs.CL TIER_1 English(EN) · Lucas Bandarkar, Alan Ansell, Trevor Cohn ·

    大型推理模型难以跨脚本迁移参数知识

    arXiv:2603.17070v2 Announce Type: replace Abstract: In this work, we analyze shortcomings in cross-lingual knowledge transfer in large, modern reasoning LLMs. We demonstrate that the perceived gap in knowledge transfer is primarily a script barrier. First, we conduct an observati…

  122. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过自适应结构化非结构化数据实现令牌高效数据推理代理

    Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this tod…

  123. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越人类监督扩展大型推理模型:通往超级智能的道路

    This work proposes a structured ladder for scaling large reasoning models beyond human supervision by tracing autonomous rewards and self-generated experience, while identifying risks and evaluation dimensions.

  124. Hugging Face Daily Papers TIER_1 English(EN) ·

    无知还是无能?为 LLM 代理构建知识门控、可验证的任务

    A protocol separates task instructions from private convention artefacts to explicitly test agent dependence on hidden knowledge, validated by calibration tasks showing near-zero performance without access.

  125. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型中的识别-拒绝不一致:为何模型会回答结构上无法回答的问题

    Large language models encode whether structurally impossible math or code prompts are unanswerable via a hidden-state direction, but fail to abstain because this recognition signal is misaligned with safety-refusal pathways, indicating a routing rather than encoding failure.

  126. arXiv cs.AI TIER_1 English(EN) · Maciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler ·

    并行与分布式推理语言模型的性能基础

    arXiv:2608.27046v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models (RL…

  127. arXiv cs.AI TIER_1 English(EN) · Sachin Gopal Wani, Ajay Dholakia, David Ellison ·

    推理税:LLM推理在不同任务类型和部署环境下的代币经济学

    arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric…

  128. arXiv cs.AI TIER_1 English(EN) · Haizhao Fan, Yuchi Xiong, Jize Wang, Xinping Guan, Xinyi Le ·

    SymbolLKG:通过逻辑知识图谱和符号求解器实现可验证的逻辑推理

    arXiv:2608.26836v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language understanding, yet they struggle with strict multi-step reasoning, frequently suffering from hallucinations and inconsistency. Existing soluti…

  129. arXiv cs.AI TIER_1 English(EN) · Zike Yuan, Han Zhang, Jianzhi Yan, Le Liu, Cai Ke, Huozhi Zhou, Jian Xie, Jiran Yin, Yukun Cao, Yue Yu, Hui Wang, Ming Liu, Bing Qin ·

    GRAIN:通过不变性奖励的代理强化学习实现真实世界图推理中的名称和叙事转换的桥梁

    arXiv:2608.27142v1 Announce Type: new Abstract: Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting…

  130. arXiv cs.AI TIER_1 English(EN) · Philipp Schr\"oppel ·

    利用以用户为中心的思维链推理改进 LLM 可解释性

    arXiv:2608.26166v1 Announce Type: cross Abstract: Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - intermediate steps toward solutions - open up high-stakes applications by enabling human inspect…

  131. arXiv cs.AI TIER_1 English(EN) · Yunxin Sun, Abulhair Saparov ·

    语言模型遵循奥卡姆剃刀吗?归纳和溯因推理中简洁性的评估

    arXiv:2509.03345v3 Announce Type: replace Abstract: Non-deductive reasoning, encompassing inductive and abductive reasoning, is essential in addressing complex real-world questions. One key feature of inductive and abductive reasoning is that there are many valid hypotheses; the …

  132. arXiv cs.AI TIER_1 English(EN) · Yuxiang Wang, Junhao Gan, Shengxiang Gao, Shenghao Ye, Zhengyi Yang, Jianzhong Qi ·

    超越线性化:用于表格推理的归因表格图

    arXiv:2601.08444v2 Announce Type: replace Abstract: Table reasoning, a task to answer questions by reasoning over data presented in tables, is an important topic due to the prevalence of knowledge stored in tabular formats. Recent solutions use Large Language Models (LLMs) for th…

  133. arXiv cs.AI TIER_1 English(EN) · Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He ·

    从准确性到鲁棒性:基于规则和模型的验证器在数学推理中的研究

    arXiv:2505.22203v3 Announce Type: replace-cross Abstract: Trustworthy verifiers are essential for the success of reinforcement learning with verifiable reward (RLVR), which is the core methodology behind various large reasoning models such as DeepSeek-R1. In complex domains like …

  134. arXiv cs.CL TIER_1 English(EN) · Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth ·

    神经符号PRM:通过结构化追踪和符号验证增强科学推理

    arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed, mathe…

  135. arXiv cs.CL TIER_1 English(EN) · Fei Ding ·

    千图假说:一种可测试的、关于存储库级代码推理中任务条件化关系物化的假说

    arXiv:2608.26602v1 Announce Type: cross Abstract: Large software repositories are often beyond model context limits. Training repository knowledge into models is costly and quickly stale, while local retrieval can miss scattered requirements, and explicit relation graphs add ongo…

  136. arXiv cs.LG TIER_1 English(EN) · Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang ·

    理解LLM推理的进化策略:比GRPO更广泛的推理覆盖范围

    arXiv:2608.27351v1 Announce Type: new Abstract: Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to …

  137. Hugging Face Daily Papers TIER_1 English(EN) ·

    理解LLM推理的演化策略:比GRPO更广泛的推理覆盖范围

    Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group …

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    GRAIN:通过不变性奖励的代理强化学习实现真实世界图推理中的名称和叙事转变的桥梁

    Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting topological structures from noisy text is highl…

  139. Hugging Face Daily Papers TIER_1 English(EN) ·

    并行与分布式推理语言模型的性能基础

    Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models (RLMs) such as DeepSeek-R1, o3, and Kimi k1.5 show th…

  140. Hugging Face Daily Papers TIER_1 English(EN) ·

    SymbolLKG:通过逻辑知识图谱和符号求解器实现可验证的逻辑推理

    Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language understanding, yet they struggle with strict multi-step reasoning, frequently suffering from hallucinations and inconsistency. Existing solutions like Chain-of-Thought (CoT) lack rigorous ve…

  141. arXiv cs.LG TIER_1 English(EN) · Pin-Han Ho, Limei Peng, Yiming Miao, Yan Jiao ·

    认知记忆:自维护智能系统的有效性层

    arXiv:2510.16899v2 Announce Type: replace Abstract: AI memory mechanisms primarily focus on preserving information content, often neglecting the validity conditions under which knowledge remains applicable, leading to semantic coordinate drift when agents move, change sensors, or…

  142. arXiv cs.LG TIER_1 English(EN) · Samuel Lippl, Thomas McGee, Kimberly Lopez, Ziwen Pan, Pierce Zhang, Salma Ziadi, Oliver Eberle, Ida Momennejad ·

    AlgoTrace:语言模型中推理的算法原语和组合几何

    arXiv:2510.15987v3 Announce Type: replace Abstract: How do inference time and latent computations enable large language models (LLMs) to solve multi-step reasoning problems? We introduce AlgoTrace, a framework for tracing and steering algorithmic operations in the model latent sp…

  143. arXiv cs.CL TIER_1 English(EN) · Chung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei Weng ·

    ReFIne:一个用于可信赖的大型推理模型的框架,具备可靠性、忠实性和可解释性

    arXiv:2510.09062v2 Announce Type: replace Abstract: Recent advances in long chain-of-thought (CoT) reasoning have largely prioritized answer accuracy and token efficiency, while overlooking aspects critical to trustworthiness. We argue that usable reasoning systems must be trustw…

  144. arXiv cs.CL TIER_1 English(EN) · Srimonti Dutta, Akshata Kishore Moharir ·

    LLM数据代理的追踪完整性:面向现实世界系统中可审计结构化推理的愿景

    arXiv:2608.26036v1 Announce Type: cross Abstract: Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability c…

  145. arXiv cs.CL TIER_1 English(EN) · Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen, Yu Yang ·

    Reflection Steering: 在激活空间中解耦反射与推理以实现令牌高效推理

    arXiv:2608.25542v1 Announce Type: cross Abstract: Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection …

  146. arXiv cs.CL TIER_1 English(EN) · Rongchen Zhao, Yu Chen, Juyuan Wang, Zhouting Mo, Jianxing Yu, Wenqing Chen, Jingping Liu ·

    PonsRAG:受Pons启发的RAG,用于协调长篇叙事推理的认知岛屿桥梁

    arXiv:2608.25486v1 Announce Type: cross Abstract: Long Narrative Reasoning is an essential capability for processing and reasoning over complex narratives. While retrieval-augmented generation provides a promising framework, existing methods still face two critical challenges: co…

  147. arXiv cs.CL TIER_1 English(EN) · Lam So, Canhui Wu, Han Lin ·

    GRIP:面向高效推理的细粒度奖励引导参数插值

    arXiv:2608.25583v1 Announce Type: new Abstract: Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference cost and latency. In contrast, instruction-tuned mod…

  148. arXiv cs.CL TIER_1 English(EN) · Jinpu Jiang, Xuan Wu, Wenhao Song, Bo Yang, You Zhou, Hongwei Ge, Heow Pueh Lee, Yanchun Liang, Chunguo Wu ·

    ReliableRAG:通过可靠性引导的推理链对抗检索增强生成中的错误信息

    arXiv:2608.25487v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in…

  149. arXiv cs.CL TIER_1 English(EN) · Casey Kennington ·

    人工智能系统计算语义学入门

    arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed abo…

  150. Hugging Face Daily Papers TIER_1 English(EN) ·

    理解LLM推理的演化策略:比GRPO更广泛的推理覆盖范围

    Evolution strategies improve reasoning diversity and Pass@K over GRPO through sparse functional updates and population diversity, supporting a hybrid training approach.

  151. Hugging Face Daily Papers TIER_1 English(EN) ·

    代码即世界:代理发现可执行世界表征以进行物理推理

    Code-as-World represents physical environments as executable code to enable quantitative reasoning and scalable supervision for vision-language models.

  152. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM数据代理的追踪完整性:面向现实世界系统中可审计结构化推理的愿景

    Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation re…

  153. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Chunguo Wu ·

    ReliableRAG:通过可靠性引导的推理链对抗检索增强生成中的错误信息

    Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in news and social media poses a serious challenge…

  154. arXiv cs.AI TIER_1 English(EN) · Zhen Bi, Xueshu Chen, Yan Wang, Zhizhi Peng, Haosen Hong, Zhen Wang, Zhixuan Chu, Bingyu Zhu, Jungang Lou ·

    记忆并非总是必需:科学推理中的条件记忆特征

    arXiv:2608.23982v1 Announce Type: new Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. Conditional memory provides an explicit lookup pathway that complements dense neural representa…

  155. arXiv cs.AI TIER_1 English(EN) · Zhenyu Wu, Siyuan Chen, Changchun Yang, Jiaqi Dong, Min Zhou, Ali Almadan, Talal Hammad, Faisal Wahbo, Aminullah Tora, Mona Alshahrani, Xin Gao ·

    TRACE:一个基于证据的大型推理模型安全评估基准

    arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchm…

  156. arXiv cs.AI TIER_1 English(EN) · Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han, Ligeng Zhu ·

    Parason:揭示大语言模型推理中的子任务与试错并行性

    arXiv:2608.24658v1 Announce Type: new Abstract: Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for …

  157. arXiv cs.AI TIER_1 English(EN) · Shaojie Zhu, Zhaobin Wang, Chengxiang Zhuo, Hui Lu, Bo Hu, Zang Li ·

    Olapa-MCoT:提升大语言模型中文数学推理能力

    arXiv:2312.17535v2 Announce Type: replace Abstract: In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking has led to large language models (LLMs) attracting widespread attention. However, restricted by expensive costs, many studies have to…

  158. arXiv cs.LG TIER_1 English(EN) · Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong, Fumiya Uchiyama, Kenji Kubo, Kohei Hayashi, Masahiro Suzuki, Yutaka Matsuo ·

    在推理时使用读出反馈引导循环推理器

    arXiv:2608.24136v1 Announce Type: new Abstract: Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more st…

  159. arXiv cs.AI TIER_1 English(EN) · Md Mahadi Hasan Nahid, Davood Rafiei ·

    PARTAB:面向可扩展表格理解的结构化证据的感知分区推理

    arXiv:2608.24082v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for r…

  160. arXiv cs.AI TIER_1 English(EN) · Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie ·

    递归式代理推理

    arXiv:2608.23956v1 Announce Type: new Abstract: Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are often evaluated in isolation, making their gains difficult to compare across models, benchmarks, and evaluation pipelines. We introdu…

  161. Hugging Face Daily Papers TIER_1 English(EN) ·

    PARTAB:用于可扩展表格理解的结构化证据的感知分区推理

    Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for reasoning. Existing approaches typically reason ove…

  162. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Davood Rafiei ·

    PARTAB:用于可扩展表格理解的感知分区结构化证据推理

    Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for reasoning. Existing approaches typically reason ove…

  163. arXiv cs.CL TIER_1 English(EN) · Murat Dura, Serkan \"Ozt\"urk, Selma Tekir ·

    通过顺序激活打补丁实现思维链推理的机制可解释性

    arXiv:2608.22332v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we i…

  164. arXiv cs.AI TIER_1 English(EN) · Yujie Zhang, Bin Gao, Tulika Mitra ·

    SAEM:用于链式思考推理中内存高效 MoE 推理的阶段感知专家管理

    arXiv:2608.21614v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity t…

  165. arXiv cs.AI TIER_1 English(EN) · Orion Powers, Daniella Seum, Khaled Slhoub ·

    更准确还是更高效?评估本地部署的紧凑型开源语言模型在数学推理方面的表现

    arXiv:2608.22048v1 Announce Type: new Abstract: Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy while fewer quantify local runtime and energy, characterize failure modes, or ap…

  166. arXiv cs.AI TIER_1 English(EN) · Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman, Afsaneh Fatemi ·

    小型推理模型是函数调用的指令遵循者

    arXiv:2608.22472v1 Announce Type: new Abstract: Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs function-calling accuracy through fine-tuning, reinforcement learning (RL), and multi-agent fra…

  167. arXiv cs.AI TIER_1 English(EN) · Min Chen, Shengjun Zhang, Yuxin Li, Zhang Zhang, Xin Fei, Chong Xia, Yueqi Duan ·

    ParallelWorld:具身推理的测试时缩放

    arXiv:2608.22971v1 Announce Type: new Abstract: Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of e…

  168. arXiv cs.AI TIER_1 English(EN) · Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu, Ziyi Wang, Xun Zhao, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen ·

    分块推理RL是否真的优于SFT?在无CoT数据下重新审视训练策略

    arXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method train…

  169. arXiv cs.AI TIER_1 English(EN) · Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang, Di Wang ·

    通过安全方向惩罚缓解推理诱导的失准

    arXiv:2608.23497v1 Announce Type: new Abstract: Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious c…

  170. arXiv cs.AI TIER_1 English(EN) · Weihang Pan, Zhengxu Yu, Yuxiang Zhang, Wenzhi Li, Zhongming Jin, Binbin Lin, Xiaofei He, Jieping Ye ·

    ChainPrune:评估和减少长链思维推理中的冗余

    arXiv:2608.21860v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoning. However, advanced Large Reasoning Models (LRMs…

  171. arXiv cs.AI TIER_1 English(EN) · Lihui Liu, Zihao Wang, Hanghang Tong ·

    知识图谱上的神经符号推理:从查询视角进行的调查

    arXiv:2412.10390v2 Announce Type: replace Abstract: Knowledge graph reasoning is pivotal in various domains such as data mining, artificial intelligence, the Web, and social sciences. These knowledge graphs function as comprehensive repositories of human knowledge, facilitating t…

  172. arXiv cs.AI TIER_1 English(EN) · Zhejian Lai, Xiang Geng, Zhijun Wang, Yang Bai, Jiahuan Li, Rongxiang Weng, Jingang Wang, Xuezhi Cao, Xunliang Cai, Shujian Huang ·

    AdaR:为大型语言模型配备自适应推理的框架

    arXiv:2510.04617v3 Announce Type: replace Abstract: Mathematical reasoning is a primary indicator of large language models (LLMs) intelligence. However, existing LLMs exhibit failures in robustness and generalization. This paper attributes these deficiencies to spurious reasoning…

  173. arXiv cs.CL TIER_1 English(EN) · Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu ·

    超越事实知识:大型语言模型中步骤级程序规则推理的基准测试与学习

    arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a large…

  174. arXiv cs.CL TIER_1 English(EN) · Yi Lu, Deyang Kong, Jianing Wang, Linsen Guo, Xue Wang, Qi Guo, Tao Gui, Xuanjing Huang, Wei Ye, Shikun Zhang, Wei Wang ·

    面向复杂推理的块扩散语言模型的自适应测试时计算分配

    arXiv:2602.09555v3 Announce Type: replace Abstract: Recent advances in block diffusion language models have demonstrated competitive performance and strong scalability on reasoning tasks. However, their test-time compute allocation remains largely unexplored, leaving a critical s…

  175. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过安全方向惩罚缓解推理诱导的失准

    Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safety of LLM reasoning. Cross-a…

  176. Hugging Face Daily Papers TIER_1 English(EN) ·

    ParallelWorld:具身推理的测试时缩放

    Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of embodied reasoning from static perception toward …

  177. arXiv cs.CL TIER_1 English(EN) · Ravisri Valluri, Tung Nguyen, Aditya Grover ·

    面向更快速推理模型的自思索

    arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor …

  178. arXiv cs.AI TIER_1 English(EN) · Haorui Xu, Yuzhou Zhu, Liyuan Gao ·

    DirEAG:用于校准数学推理中口头化置信度的狄利克雷证据聚合

    arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steerin…

  179. arXiv cs.AI TIER_1 English(EN) · Said Slaoui ·

    S-AI-Recursive:收敛递归推理

    arXiv:2605.13872v2 Announce Type: replace-cross Abstract: This article introduces S-AI-Recursive, a bio-inspired Sparse Artificial Intelligence architecture in which reasoning is implemented as a hormonally regulated closed-loop iteration rather than a single feed-forward pass. T…

  180. arXiv cs.CL TIER_1 English(EN) · Lan Zhang, Marco Valentino, Jordan Meadows, Andre Freitas ·

    超越黄金标准:用于形式数学推理的语言模型评委的认识论集成

    arXiv:2506.10903v2 Announce Type: replace Abstract: Statement autoformalization plays a crucial role in formal mathematical reasoning by enabling the automatic translation of natural language statements into formal languages. While recent advances using large language models (LLM…

  181. arXiv cs.CL TIER_1 English(EN) · Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang, Tingwen Liu ·

    内存增强解锁高效思维链推理

    arXiv:2608.21265v1 Announce Type: new Abstract: Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may …

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    分块推理RL是否真的优于SFT?在无CoT数据下重新审视训练策略

    Mixed supervised fine-tuning on combined reasoning corpora outperforms next-chunk reinforcement learning in efficiency and final accuracy across mathematical and out-of-domain tasks.

  183. arXiv cs.CL TIER_1 English(EN) · Philipp Hellwig, Willem Zuidema, Claire E. Stevenson, Martha Lewis ·

    Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning

    arXiv:2604.06501v2 Announce Type: replace-cross Abstract: Analogical reasoning is a hallmark of human intelligence, enabling us to solve new problems by transferring knowledge from one situation to another. Yet, developing artificial intelligence systems capable of robust human-l…

  184. arXiv cs.LG TIER_1 English(EN) · Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi ·

    Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    arXiv:2608.19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these…

  185. arXiv cs.AI TIER_1 English(EN) · Muhan Gao, Zih-Ching Chen, Kuan-Hao Huang ·

    第一滴墨水:长上下文推理中干扰信息的非线性影响

    arXiv:2605.10828v2 Announce Type: replace Abstract: As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understanding how distracting information affects long-context performance becomes criti…

  186. arXiv cs.AI TIER_1 English(EN) · Gijs Kassenaar, Zhao Yang, Vincent Fran\c{c}ois-Lavet ·

    学习何时思考:用于测试时间计算分配的自适应推理

    arXiv:2608.20256v1 Announce Type: new Abstract: Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on d…

  187. arXiv cs.AI TIER_1 English(EN) · Yiting Qu, Ziqing Yang, Chi Cui, Ye Leng, Junjie Chu, Yang Zhang ·

    EchoCoT:从大型推理模型中提取隐藏的思维链

    arXiv:2608.20055v1 Announce Type: cross Abstract: Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be directly extracted from black-box models remains large…

  188. arXiv cs.AI TIER_1 English(EN) · Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao ·

    迈向通用具身智能:整合大型语言模型、知识库和推理能力,构建下一代AI代理

    arXiv:2608.19794v1 Announce Type: new Abstract: The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI). This paper reviews the evolution of LLM-center…

  189. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward fee…

  190. arXiv cs.AI TIER_1 English(EN) · Pratik Ghawate ·

    FinRCA-Bench:金融AI系统的证据检索与推理基准测试

    arXiv:2608.18534v1 Announce Type: new Abstract: Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagno…

  191. arXiv cs.LG TIER_1 English(EN) · Lirui Luo, Guoxi Zhang, Hongming Xu, Rongqing Li, Cong Fang, Lifeng Fan ·

    Continual Reasoning Gym: 诊断和利用持续RLVR中的共享推理

    arXiv:2608.18574v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study contin…

  192. arXiv cs.CL TIER_1 English(EN) · Yajie Yin ·

    评分评分者:LLM推理的验证自主性等级(L0-L5)

    arXiv:2608.19009v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literatur…

  193. arXiv cs.AI TIER_1 English(EN) · Gilad Abiri, Emanuel V. Towfigh ·

    认识论上的从属:生成式AI与知识基础设施

    arXiv:2608.18758v1 Announce Type: cross Abstract: Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full bread…

  194. arXiv cs.AI TIER_1 English(EN) · Zuocheng Ying, Yang Yang, Yumou Wu, Chuanbo Zhu, Jiarui Wang, Ziqi Wu, Jingming Cai, Junqing Yu, Zikai Song ·

    从存储到访问:通过显式提示和隐式推理实现大语言模型参数化知识的可验证激活

    arXiv:2608.18581v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. Existing end-to-end methods entangle …

  195. arXiv cs.AI TIER_1 English(EN) · Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang ·

    LLaMA 3.1 8B 结构感知数值推理的机制可解释性

    arXiv:2608.18419v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with w…

  196. arXiv cs.AI TIER_1 English(EN) · Zijie Meng, Xiwei Dai, Yixuan Tang, Jin Hao, Yang Feng, Fudong Zhu, Xiaoqiang Liu, Shaosheng Cao, Zuozhu Liu ·

    DentAgent:面向多模态牙科推理的以证据为中心的多代理协调

    arXiv:2608.18878v1 Announce Type: new Abstract: Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D den…

  197. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zuozhu Liu ·

    DentAgent:面向多模态牙科推理的以证据为中心的多代理协调

    Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D dental data. Most existing dental AI systems remain…

  198. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Pratik Ghawate ·

    FinRCA-Bench:金融AI系统的证据检索与推理基准测试

    Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase ord…

  199. Hugging Face Daily Papers TIER_1 English(EN) ·

    FinRCA-Bench:为金融AI系统进行证据检索和推理的基准测试

    Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase ord…

  200. arXiv cs.CL TIER_1 English(EN) · Eduardo S\'anchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer ·

    IOL-AI 挑战赛:一项旨在推进语言推理的开放性挑战

    arXiv:2608.18011v1 Announce Type: new Abstract: Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-A…

  201. arXiv cs.AI TIER_1 English(EN) · Yeabin Moon ·

    思考的代价:推理成本作为模型特定的API契约

    arXiv:2608.16956v1 Announce Type: new Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reason…

  202. arXiv cs.AI TIER_1 English(EN) · Xing Wei, Changmeng Zheng, XiaoYong Wei, Xiufen Ye, Qing Li ·

    DeAR:通过能力基础和协作思维导航实现去中心化代理推理

    arXiv:2608.17282v1 Announce Type: new Abstract: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentrali…

  203. arXiv cs.AI TIER_1 English(EN) · Guozheng Sun ·

    SignalReasoner:评估3B模型在信号数学推理中的上限

    arXiv:2608.17301v1 Announce Type: new Abstract: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their applica…

  204. arXiv cs.AI TIER_1 English(EN) · Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li ·

    Co-RL:多智能体强化学习中多样化群体涌现无监督推理

    arXiv:2608.17253v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward).…

  205. Hugging Face Daily Papers TIER_1 English(EN) ·

    Co-RL:多智能体强化学习中多样化群体涌现无监督推理

    Co-RL enables unsupervised reasoning via cooperative multi-agent reinforcement learning with peer-derived rewards, improving performance across text and vision tasks without ground-truth labels.

  206. arXiv cs.AI TIER_1 English(EN) · Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu ·

    重新思考SFT推理中的泛化:对优化、数据和模型能力的条件分析

    arXiv:2604.06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and…

  207. arXiv cs.CL TIER_1 English(EN) · Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li ·

    $R^3$-Bench:LLM在共享预算下难以进行资源理性推理

    arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not…

  208. arXiv cs.LG TIER_1 English(EN) · Jules Soria, Alban Grastien, Romain Xu-Darme, Julien Girard-Satabin, Zakaria Chihani, Daniela Cancila ·

    超越 $L_2$:将溯因潜在解释推广到多样化的基于原型的架构

    arXiv:2608.16773v1 Announce Type: new Abstract: Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic…

  209. arXiv cs.CL TIER_1 English(EN) · Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li ·

    提问、约束或弃权:用于缺失前提推理的强化学习

    arXiv:2608.16554v1 Announce Type: new Abstract: Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the…

  210. arXiv cs.AI TIER_1 English(EN) · Lirui Teng ·

    GRIP:通过信息受限前提进行接地推理

    arXiv:2608.16776v1 Announce Type: new Abstract: High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence functionally irrelevant. We call this failure mode query dominance. To address it, we introduce \…

  211. arXiv cs.AI TIER_1 Español(ES) · Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen ·

    ParaTempo:通过时间置信度实现高效并行推理

    arXiv:2608.16425v1 Announce Type: new Abstract: Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these para…

  212. arXiv cs.AI TIER_1 English(EN) · Bo Wen, Yuhao Chen, Erhan Bilal, Carla Agurto Rios, Chen Wang, Junchen Jiang ·

    发散-收敛推理:通过结构化解决方案合成扩展测试时计算

    arXiv:2608.15303v1 Announce Type: new Abstract: Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase pri…

  213. arXiv cs.AI TIER_1 English(EN) · Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang ·

    未被书写的基准:多模态机器学习在抽象感知推理方面的新挑战

    arXiv:2608.14558v1 Announce Type: new Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative p…

  214. arXiv cs.AI TIER_1 English(EN) · Shufeng Kong, Xiaochuan Zhang, Caihua Liu ·

    职位:神经约束推理的认证正确性需要符号集成

    arXiv:2608.14569v1 Announce Type: new Abstract: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the mode…

  215. Hugging Face Daily Papers TIER_1 English(EN) ·

    SignalReasoner:评估3B模型在信号数学推理中的上限

    Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relat…

  216. Hugging Face Daily Papers TIER_1 Español(ES) ·

    ParaTempo:通过时间置信度实现高效并行推理

    ParaTempo improves parallel reasoning efficiency by using temporal confidence to dynamically prune, retire, and reallocate reasoning branches without synchronization.

  217. arXiv cs.AI TIER_1 English(EN) · Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn, Soheil Feizi ·

    Early Stopping for Large Reasoning Models via Confidence Dynamics

    arXiv:2604.04930v2 Announce Type: replace-cross Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key chal…

  218. arXiv cs.AI TIER_1 English(EN) · Feng Xiong, Leyan Xue, Hongyu Lin ·

    纠正无法看到的内容:多模态推理器中感知蒸馏的信用分配

    arXiv:2607.28336v3 Announce Type: replace Abstract: On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subsequent reasoning. Perception Success Rate (PSR), e…

  219. arXiv cs.AI TIER_1 English(EN) · Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen ·

    元认知瓶颈:日本谜题揭示了推理AI在洞察力和自我评估方面的根本性局限

    arXiv:2509.14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.…

  220. arXiv cs.AI TIER_1 English(EN) · Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig ·

    放大不等于预测:思考模型中的推理行为

    arXiv:2608.13760v1 Announce Type: cross Abstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces loo…

  221. arXiv cs.AI TIER_1 English(EN) · Cuong Dang, Hoang Anh Just, Ruoxi Jia ·

    容量依赖的数据选择对推理的影响

    arXiv:2608.13721v1 Announce Type: cross Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that …

  222. arXiv cs.AI TIER_1 English(EN) · Dayuan Zhao, Shengcao Cao, Yu-Xiong Wang, Liang-Yan Gui ·

    在潜在空间思考,用语言解释:自解释潜在推理

    arXiv:2608.13570v1 Announce Type: cross Abstract: Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing rea…

  223. arXiv cs.AI TIER_1 English(EN) · Zhelun (Allen), Wu ·

    从不数字:AI系统答案被视为事实时的结构性弃权

    arXiv:2608.13926v1 Announce Type: new Abstract: Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent …

  224. arXiv cs.AI TIER_1 English(EN) · Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz ·

    BiasTrace:将推理行为与大型语言模型中的偏见输出联系起来

    arXiv:2608.14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and mitigating bias, it largely focuses on final outputs…

  225. arXiv cs.AI TIER_1 English(EN) · Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi, David Lo ·

    Second Thought:LLM代理并行推理,行动与观察

    arXiv:2608.13667v1 Announce Type: new Abstract: LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is confined to the Thought phase: while the agent serializes an action and waits for the environment, its reasoning is fro…

  226. arXiv cs.AI TIER_1 English(EN) · Panjing He, Mingyue Cheng, Yucong Luo, Li Li, Xiaohan Zhang ·

    SheetCompass:用于 Agentic 电子表格推理的分层关系图

    arXiv:2608.14452v1 Announce Type: new Abstract: Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet reasoning remains challenging for large language models (LLMs). Real-world workbooks often contain implicit cross-tabl…

  227. Hugging Face Daily Papers TIER_1 English(EN) ·

    $R^3$-Bench:LLM在共享预算下难以进行资源理性推理

    In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same mo…

  228. Hugging Face Daily Papers TIER_1 English(EN) ·

    R^3-Bench:LLM在共享预算下进行资源理性推理时面临困难

    R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks.

  229. arXiv cs.AI TIER_1 English(EN) · Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin ·

    你是在跟我讲逻辑吗?评估语言模型的三段论推理能力

    arXiv:2608.12374v1 Announce Type: cross Abstract: Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation…

  230. arXiv cs.AI TIER_1 English(EN) · Rachel Lawrence, Jacqueline Maasch ·

    职位:推理是一个可学习的基于规则的过程

    arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite im…

  231. arXiv cs.AI TIER_1 English(EN) · Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar ·

    推理陪审团:多模型共识用于评估推理痕迹

    arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning beh…

  232. arXiv cs.AI TIER_1 English(EN) · Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini ·

    严苛的法官:在沉默、压力和坚持下的认知稳定性

    arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-pro…

  233. arXiv cs.AI TIER_1 English(EN) · Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang ·

    ReflectFact:用于改进多跳事实验证中理解和推理的自反思代理

    arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent …

  234. arXiv cs.AI TIER_1 English(EN) · Yubo Li, Ramayya Krishnan, Rema Padman ·

    大型语言模型知道约束但未使用:实用约束推理中的激活瓶颈

    arXiv:2608.12321v1 Announce Type: cross Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditiona…

  235. arXiv cs.CL TIER_1 English(EN) · Gaurav Najpande, Tampu Ravi Kumar, Manan Roy Choudhury, Neha Valeti, Yanjie Fu, Vivek Gupta ·

    QUIETT:查询无关的表格转换,实现鲁棒推理

    arXiv:2602.20017v2 Announce Type: replace Abstract: Real-world tables often contain schema inconsistencies, heterogeneous value formats, and implicit relational structures that degrade table reasoning and question answering. Existing approaches address these issues at query time,…

  236. arXiv cs.AI TIER_1 English(EN) · Obed Junias, Maria Leonor Pacheco ·

    从原子证据到逻辑组合:对复合答案选项的结构化组合推理

    arXiv:2608.12836v1 Announce Type: cross Abstract: Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEI…

  237. Hugging Face Daily Papers TIER_1 English(EN) ·

    永不为数字:AI系统答案被视为事实时的结构性弃权

    Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indistinguishable at the point of …

  238. Hugging Face Daily Papers TIER_1 English(EN) ·

    Second Thought:LLM 代理并行推理,行动与观察

    Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy.

  239. arXiv cs.AI TIER_1 English(EN) · Atahan Dokme, Benjamin Reichman, Larry Heck ·

    TEMPER:量化推理中的情绪扰动测试

    arXiv:2604.07801v2 Announce Type: replace-cross Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does emot…

  240. arXiv cs.AI TIER_1 English(EN) · Valentin Rodionov, Shamil Assylbekov ·

    TRACES:大型语言模型科学推理中的认知可靠性基准

    arXiv:2608.11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literatur…

  241. arXiv cs.AI TIER_1 English(EN) · Suyash Mishra ·

    本地验证无法检测非可移植性:一种关于智能体推理中上下文保持的同调理论

    arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that…

  242. arXiv cs.AI TIER_1 English(EN) · Yoshinori Watanabe ·

    脱离支持的障碍:为何语义安全约束不是学习问题的不变性,以及这对先验设计、约束和验证意味着什么

    arXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data di…

  243. Hugging Face Daily Papers TIER_1 English(EN) ·

    从原子证据到逻辑组合:化合物选项上的结构化组合推理

    A framework decomposes compound logical options into atomic judgments and uses constrained optimization to improve reasoning over AND, OR, and NEITHER/NOR operators.

  244. Hugging Face Daily Papers TIER_1 English(EN) ·

    放大不等于预测:思考模型中的推理行为

    Reasoning training amplifies deliberative behaviors like self-correction more than high-correctness behaviors such as confidence calibration, revealing a gap between amplified and correctness-linked reasoning patterns.

  245. arXiv cs.AI TIER_1 English(EN) · Rose Niousha, Minwoo Kang, Narges Norouzi ·

    深入学生思维:联合建模LLM学生模拟器中的潜在推理与行动

    arXiv:2608.10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as ev…

  246. arXiv cs.AI TIER_1 English(EN) · Si'an Xie (Beijing University of Posts and Telecommunications), Jiaxun Liu (Peking University), Biao Yang (Kuaishou Technology), Wei Yuan (Kuaishou Technology), Fan Yang (Kuaishou Technology), Tingting Gao (Kuaishou Technology), Ming Wu (Beijing Universi… ·

    从推理深度到推理广度:评估大型语言模型的多点联想推理能力

    arXiv:2608.10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unex…

  247. arXiv cs.AI TIER_1 English(EN) · Tughanbulut Kurtulush ·

    思维链何时有益何时有害:LLM推理串行深度瓶颈的实证研究

    arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds o…

  248. arXiv cs.AI TIER_1 English(EN) · Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat, Soumyabrata Pal, Ramasuri Narayanam, Dinesh Manocha ·

    ThinkRetrieve:用于测试时扩展的检索增强推理轨迹

    arXiv:2608.10928v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminis…

  249. arXiv cs.AI TIER_1 English(EN) · Xin Xu ·

    推理捷径与价值对称:对称性允许什么,架构实现什么,优化选择什么

    arXiv:2608.10420v1 Announce Type: new Abstract: Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of valu…

  250. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Shamil Assylbekov ·

    TRACES:大型语言模型科学推理中的认知可靠性基准

    Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literature, a capability that has not yet been directly mea…

  251. arXiv cs.AI TIER_1 English(EN) · Minhan Cho, Jimin Kweon ·

    复现和压力测试两种LLM推理可靠性方法:测试时概率聚合与逻辑表示编辑

    arXiv:2608.08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models…

  252. arXiv cs.LG TIER_1 English(EN) · Zachary Horvitz, Raghav Singhal, Hao Zou, Carles Domingo-Enrich, Zhou Yu, Rajesh Ranganath, Kathleen McKeown ·

    使用 MDLMs 重新思考推理:提前退出、事后推理及其他

    arXiv:2510.19990v2 Announce Type: replace Abstract: The reasoning paradigm, where language models reason before answering, has enabled breakthroughs on tasks such as mathematical problem-solving. While current tooling for reasoning is built around next-token prediction trained mo…

  253. arXiv cs.CL TIER_1 English(EN) · Hexuan Wang, Yaxuan Ren, Srikar Bommireddypalli, Shuxian Chen, Adarsh Prabhudesai, Rongkun Zhou, Elina Baral, Philipp Koehn ·

    SciTaRC:一个用于语言推理和复杂计算的带计划注释的科学表格问答基准

    arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To enable fine-grained diagnostic analysis beyond end-task accuracy, SciTaRC pairs ea…

  254. arXiv cs.CL TIER_1 English(EN) · Tianjun Zhong, Linyang He, Ziyang Li, Nima Mesgarani ·

    从链式到有向无环图:探究大型语言模型推理的图结构

    arXiv:2601.17593v3 Announce Type: replace Abstract: Recent progress in large language models has renewed interest in how multi-step reasoning is represented internally. While prior work often treats reasoning as a linear chain, many reasoning problems can be more naturally modele…

  255. arXiv cs.CL TIER_1 English(EN) · Mengxi Xiao, Kailai Yang, Pengde Zhao, Enze Zhang, Ziyan Kuang, Zhiwei Liu, Weiguang Han, Shu Liao, Lianting Huang, Guojun Xiong, Victor Gutierrez Basulto, Jinpeng Hu, Min Peng, Qianqian Xie, Sophia Ananiadou ·

    MiraMind:超越答案准确性,对可靠心理健康推理进行基准测试

    arXiv:2512.09636v3 Announce Type: replace Abstract: Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into interpretations, decisions, or claims whose spec…

  256. arXiv cs.AI TIER_1 English(EN) · Hongyi James Cai, Junlin Wang, Xiaoyin Chen, Bhuwan Dhingra ·

    多少回溯才算足够?探索SFT和RL在增强LLM推理中的相互作用

    arXiv:2505.24273v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yielding significant improvements on challenging reasoning tasks through extended chains of…

  257. arXiv cs.AI TIER_1 English(EN) · Muhammad Ali Shafique, Kelly Marchisio ·

    推理大模型中的隐藏语言一致性现象

    arXiv:2608.08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingua…

  258. arXiv cs.AI TIER_1 English(EN) · Zhengze Huang, Luyang Yu, Di Hong, Xinzhe Huang, Wanyu Lin, Zhixuan Chu, Zhan Qin, Tianhang Zheng ·

    REIN:通过反思和弃权对齐弥合推理与可靠性之间的差距

    arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where fl…

  259. arXiv cs.AI TIER_1 English(EN) · Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi ·

    深思熟虑,而非聪明:推理模型未能跨问题均摊测试时间计算

    arXiv:2608.07968v1 Announce Type: cross Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency cons…

  260. arXiv cs.AI TIER_1 English(EN) · Juncheng Dong, Ding Tong, Ishan Gupta, Yuyan Wang ·

    LLM在主观任务上的推理:失败模式、缓解和动态推理路由

    arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality…

  261. Hugging Face Daily Papers TIER_1 English(EN) ·

    Thought-Level Beam Search for Reasoning

    Gambit improves reasoning model efficiency by using thought-level beam search to dynamically allocate compute to promising reasoning traces under fixed hardware budgets.

  262. Hugging Face Daily Papers TIER_1 English(EN) ·

    StructReward:用于自纠正多模态推理的高效结构化过程奖励

    Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, thereby discarding the supervisi…

  263. LessWrong (AI tag) TIER_1 English(EN) · star2vec ·

    潜在推理未见回退迹象:最终答案已然落定

    <img alt="fig0_banner_dense.png" src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789247969/lexical_client_uploads/nbvufluvrxky0i10mb5d.png" /><p><span style="white-space: pre-wrap;">Solving a hard math problem is not linear, it's trial and error.</span></p><p><span s…

  264. LessWrong (AI tag) TIER_1 English(EN) · Chi Nguyen ·

    改善概念推理能力的有利观点

    <p><span style="white-space: pre-wrap;">Below are some informal perspectives on improving AIs’ conceptual reasoning capabilities that motivate work in the area. Conceptual reasoning here means reasoning about questions where we cannot verify the answer, also don’t have data on un…

  265. arXiv cs.CV TIER_1 English(EN) · Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu ·

    代码即世界:代理发现可执行世界表征以进行物理推理

    arXiv:2608.27549v1 Announce Type: new Abstract: Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit represent…

  266. LessWrong (AI tag) TIER_1 English(EN) · Shahriar Golchin ·

    人工智能模型想被监控吗?衡量大型推理模型的可监控性倾向

    <p><span>As AI models take on increasingly high-stakes responsibilities, understanding what a model is actually doing instead of simply constraining its outputs has become one of the most important safety challenges in AI.&nbsp;</span><a href="https://arxiv.org/abs/2512.18311?ref…

  267. arXiv cs.CV TIER_1 English(EN) · Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap ·

    Investigating Relational Reasoning in VLMs

    arXiv:2608.23518v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we…

  268. arXiv cs.CV TIER_1 English(EN) · Zesheng Yang, Lingling Zhang, Xinyu Zhang, Cheng Zhang, Pengyu Li, Heng Wang, Lin Wu ·

    GLaQ:在视觉证据中进行基础潜在查询以实现多模态推理

    arXiv:2608.15517v1 Announce Type: new Abstract: Chain-of-thought reasoning has substantially improved the problem-solving capabilities of multimodal large language models. Fine-grained visual evidence, however, remains difficult to preserve and reuse across text-based reasoning s…

  269. arXiv cs.CV TIER_1 English(EN) · Xiaohan Zhang, Feng Gu, Xudong Rao, Xuhao Pan, Tao Wei, Zhou Pan, Kun Zhan ·

    ChainSpace:空间智能的链式推理范式

    arXiv:2608.15788v1 Announce Type: new Abstract: Spatial intelligence requires foundation models to maintain coherent spatial state across interactions with the physical world. However, existing data-centric approaches typically treat spatial reasoning as independent question-answ…

  270. LessWrong (AI tag) TIER_1 English(EN) · Chi Nguyen ·

    推出概念推理指数

    <p><span>Associated announcement tweet.</span></p><p><span>We are planning to release blog posts properly arguing the case for this kind of work in the future.</span></p><h1><span>tl;dr</span></h1><p><span>A core hope for managing AI risks is that AIs will help us understand the …

  271. HN — anthropic stories TIER_1 English(EN) · optimalsolver ·

    Anthropic:推出概念推理指数

  272. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    构建注重推理的大语言模型:SupraLabs推理语料库的流式传输、策展和微调实用指南

    <p>This tutorial provides a complete workflow for building a compact, reasoning-focused language model. By streaming the SupraLabs reasoning corpus from Hugging Face, we apply quality filters and curate data for Supervised Fine-Tuning (SFT). Using SmolLM2-135M-Instruct and LoRA, …

  273. Medium — fine-tuning tag TIER_1 English(EN) · Sowmithdurusoju ·

    使用 GRPO 微调创建推理模型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sowmithdurusoju/creating-reasoning-models-with-grpo-finetuning-20a36b0e57a9?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1200/1*52togsagBUurrab71b6N4A.gif" widt…

  274. Medium — fine-tuning tag TIER_1 English(EN) · Omokemiayodimeji ·

    Hop-1:将一步推理浓缩到拥有2.7亿参数的模型中

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@omokemiayodimeji/hop-1-distilling-one-reasoning-step-into-a-270-million-parameter-model-189aa724e545?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1600/1*t7YE56T…

  275. Towards AI TIER_1 English(EN) · Vedant Pandhare ·

    RLHF:从人类反馈到可验证真相:现代大型语言模型如何真正学会推理

    <p><em>14 min read · AI Research · LLM Training</em></p><p>I want to tell you about the moment I stopped thinking about AI as a prediction machine.</p><p>It happened when I was reading through the DeepSeek-R1 paper at an ungodly hour — the kind of reading you do when something ge…

  276. Medium — fine-tuning tag TIER_1 English(EN) · Okan Yenigün ·

    使用教师生成的推理微调 Qwen3–4B

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/fine-tuning-qwen3-4b-with-teacher-generated-reasoning-01d48c05e852?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2600/1*sAOLPgpTmbU8oKoRH9cplw.jpeg" width…

  277. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    首个 Token 至关重要:理解大型推理模型中的安全崩溃

    <h1> First Token Matters: Understanding Safety Collapse in Large Reasoning Models </h1> <p>The recent shift toward Large Reasoning Models (LRMs) — models that explicitly "think" through a Chain-of-Thought (CoT) before providing an answer — has promised a new era of complex proble…

  278. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    大型语言模型知识-推理权衡:为何2026年最佳模型故意减少事实——且速度前所未有

    <h2> Table of Contents </h2> <ol> <li>The Model That Spent 21 Minutes Drawing a Circle</li> <li>The LLM Knowledge-Reasoning Tradeoff: The Physics Behind It</li> <li>Why Reasoning Compresses Better Than Facts</li> <li>The Hallucination Paradox: World-Class Reasoning, High Factual …

  279. dev.to — LLM tag TIER_1 English(EN) · mayankpallai ·

    大型语言模型是如何学会推理的:SFT --> RLHF --> RLVR

    <h2> 1. The Starting Line: The Last Non-Reasoning Flagships </h2> <p>GPT-4.5, DeepSeek-V3, and Claude 3.5 Sonnet share something that has nothing to do with benchmark scores: they were the last major models built entirely on the "pretrain, then instruct-tune" recipe. All internal…

  280. dev.to — LLM tag TIER_1 English(EN) · Patience Sitati ·

    当 RAG 失效时:现代 LLM 中的锚定偏差与元认知

    <p><em>## How a bizarre Claude Code CLI bug exposed the limits of multi-turn troubleshooting loops, and why the future of AI relies on cognitive routing.</em></p> <p>I recently ran into an issue with Claude Code on Windows.</p> <p>I gave Claude everything it needed: the terminal …

  281. r/MachineLearning TIER_1 English(EN) · /u/Typical-Scene-5794 ·

    2026年潜在推理格局:BDH-CQ、HRM/TRM、Coconut [D] 路线图

    <!-- SC_OFF --><div class="md"><p>After following various arXiv papers and researcher discussions on X/bluesky about latent reasoning and continual learning, one idea which resonates strongly is that path forward (towards AGI) may depend less on generating ever-longer chains of t…

  282. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Prime Agent:通过递归子代理和持久化 REPLs 实现长时程推理的规模化

    <h1> Prime Agent: Scaling Long-Horizon Reasoning via Recursive Subagents and Persistent REPLs </h1> <p>The evolution of agentic AI has reached a critical bottleneck. While frontier models like GPT-4 and Claude 3.5 Sonnet have demonstrated impressive reasoning capabilities, their …

  283. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Next-Chunk Reasoning:为什么 RL 可能无法在无 CoT 数据上真正击败 SFT

    <h1> Next-Chunk Reasoning: Why RL Might Not Actually Beat SFT for no-CoT Data </h1> <p>The transition from standard Supervised Fine-Tuning (SFT) to Reinforcement Learning (RL) has been hailed as the primary catalyst for the reasoning capabilities observed in recent frontier model…

  284. r/LocalLLaMA TIER_1 English(EN) · /u/devildip ·

    Qwen 3.5 4B IQ2_XS:张量级分配带来+16.67%的推理性能提升

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vvc6pw/qwen_35_4b_iq2_xs_1667_reasoning_performance_from/"> <img alt="Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation" src="https://preview.redd.it/69geg9n38xkh1.png?width=140&a…

  285. r/MachineLearning TIER_1 English(EN) · /u/aglet_factorial ·

    机器学习中的认识论智能 Neurips 研讨会页面限制?[D]

    <!-- SC_OFF --><div class="md"><p>I'm aiming to submit a paper to The 3rd Workshop on Epistemic Intelligence in Machine Learning at Neurips<br /> <a href="https://eiml.cc/">https://eiml.cc/</a></p> <p>I can't find a page limit anywhere on their website and I've emailed the organi…

  286. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    TR-GRPO通过调节Token梯度稳定推理训练

    <h1> TR-GRPO Stabilizes Reasoning Training by Regulating Token Gradients </h1> <p><em>Billing Support — August 21, 2026</em></p> <h2> The failure: unlikely tokens can control the update </h2> <p>Group Relative Policy Optimization, or GRPO, is a critic-free reinforcement-learning …

  287. r/LocalLLaMA TIER_1 English(EN) · /u/SandyL925 ·

    SenseNova U1.5-Lite 全面发布:专家训练、OPD蒸馏、推理只需一个模型

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vu5dzi/sensenova_u15lite_full_release_expert_training/"> <img alt="SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference" src="https://preview.redd.it/lenqskwdhnkh1.png?w…

  288. dev.to — LLM tag TIER_1 English(EN) · Arisyn ·

    超越SQL准确性:为AI数据代理构建证据链

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj6adfu5xhoaud2tgmnnl.png"><img alt=" " height="512" …

  289. dev.to — LLM tag TIER_1 English(EN) · Ken W Alger ·

    推理账本:记住决策,而不仅仅是数据

    <p><em>Part 4 of the Building the AI Memory Stack series</em></p> <p>After finishing the previous article, I looked at the repository a little differently. The specifications were still there. The Architecture Decision Records were still there. The glossary entries were still the…

  290. r/MachineLearning TIER_1 English(EN) · /u/moschles ·

    BDH-CQ:具有循环潜在推理的上下文学习 [R]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1vov5r5/bdhcq_incontext_learning_with_recurrent_latent/"> <img alt="BDH-CQ: IN-CONTEXT LEARNING WITH RECURRENT LATENT REASONING [R]" src="https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnCYTqbJkElIQ…

  291. r/Anthropic TIER_1 English(EN) · /u/Advanced-Cat9927 ·

    超越提示工程:构建可审计的认知状态机

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1w4s6t4/beyond_prompt_engineering_building_an_auditable/"> <img alt="Beyond Prompt Engineering: Building an Auditable Epistemic State Machine" src="https://external-preview.redd.it/4lHjeGaPeUUlkP4OWfiT2s4yOpK3k…

  292. Mastodon — mastodon.social TIER_1 Русский(RU) · [email protected] ·

    研究表明,“推理模型”中的中间令牌(ITG)不能被可靠地解释为“思维痕迹”。尽管使用中间令牌进行训练有所改进

    Показывается, что промежуточные токены (ITG) в «моделях рассуждений» нельзя надежно трактовать как «следы мыслей». Хотя обучение с промежуточными токенами повышает точность, семантической связи между содержимым трасс и реальными логическими вычислениями модели часто нет. В задача…

  293. Mastodon — mastodon.social TIER_1 English(EN) · lucashendren ·

    对“只需增加更多推理”的反应的良好现实检验。在IOL语言推理挑战赛中,思维链使模型表现更差,并且约束

    Nice reality check for the 'just add more reasoning' reflex. On the IOL linguistic reasoning challenge, chain-of-thought made models perform worse, and constraining output to JSON was the single most damaging choice. The failure mode isn't lack of steps, it's that formatting and …

  294. r/OpenAI TIER_2 English(EN) · /u/Advanced-Cat9927 ·

    超越提示工程:构建可审计的认知状态机

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1w4s7a0/beyond_prompt_engineering_building_an_auditable/"> <img alt="Beyond Prompt Engineering: Building an Auditable Epistemic State Machine" src="https://external-preview.redd.it/4lHjeGaPeUUlkP4OWfiT2s4yOpK3kbu9…

  295. r/ClaudeAI TIER_2 English(EN) · /u/Enough-Plantain2785 ·

    Anthropic:推出概念推理指数

    <!-- SC_OFF --><div class="md"><p><a href="https://alignment.anthropic.com/2026/conceptual-reasoning-index/">https://alignment.anthropic.com/2026/conceptual-reasoning-index/</a></p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://www.reddit.com/user/Enough-Plantain…

  296. r/singularity TIER_2 English(EN) · /u/EducationalCicada ·

    Anthropic:推出概念推理指数

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/EducationalCicada"> /u/EducationalCicada </a> <br /> <span><a href="https://alignment.anthropic.com/2026/conceptual-reasoning-index/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/singularity/comments…