PulseAugur
EN
LIVE 17:41:10

New research tackles LLM reasoning, efficiency, and distillation challenges · 10 sources tracked

New research explores methods to improve the reasoning capabilities and efficiency of large language models (LLMs). One paper introduces "Trace as State" to enhance long-context reasoning by placing reasoning traces before the context block, showing significant performance gains on various models and datasets. Another study proposes "SALA" for semantic-aware logical alignment in in-context learning, improving demonstration selection for complex reasoning tasks. Additionally, research on "HEAL" aims to distill reasoning from LLMs into smaller models by addressing limitations in current distillation techniques, drawing inspiration from educational theory. AI

IMPACT These research efforts aim to enhance LLM reasoning accuracy, efficiency, and alignment, potentially leading to more capable and reliable AI systems.

RANK_REASON Cluster consists of multiple academic papers published on arXiv, detailing novel methods and analyses for improving LLM reasoning and evaluation.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 296 sources. How we write summaries →

New research tackles LLM reasoning, efficiency, and distillation challenges · 10 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Cluster consists of multiple academic papers published on arXiv, detailing novel methods and analyses for improving LLM reasoning and evaluation.
Source corroboration
296 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
48 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+94 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [296]

  1. arXiv cs.CL TIER_1 English(EN) · Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong ·

    Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

    arXiv:2609.19156v1 Announce Type: new Abstract: Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajector…

  2. arXiv cs.AI TIER_1 English(EN) · Chengwen Qi, Deheng Ye, Yatao Bian ·

    What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis

    arXiv:2609.19212v1 Announce Type: new Abstract: Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on si…

  3. arXiv cs.AI TIER_1 English(EN) · Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak ·

    When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

    arXiv:2609.19671v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rig…

  4. arXiv cs.AI TIER_1 English(EN) · Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar, Zachary Ives ·

    Efficiently Linking Unstructured Data for Multi-step Reasoning

    arXiv:2609.19491v1 Announce Type: cross Abstract: Modern LLMs and AI agents increasingly support data engineering workflows that integrate evidence from unstructured sources. Such pipelines typically do data retrieval, integration, and ranking before proceeding to more complex ag…

  5. arXiv cs.AI TIER_1 English(EN) · Srijith Ravikumar ·

    The Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention

    arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymanov et al. (2026) show that under safety-constrained generatio…

  6. arXiv cs.AI TIER_1 English(EN) · Anish Sathyanarayanan, Aditya Nagarsekar, Aarush Rathore ·

    Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models

    arXiv:2602.03994v3 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) prompting is widely used as a reasoning aid and is often treated as a transparency mechanism. Yet behavioral gains under CoT do not imply that the model's internal computation causally depends on the…

  7. arXiv cs.CL TIER_1 English(EN) · Magnus Boman ·

    Understanding LLM Failures: A Multi-Tape Turing Machine Analysis of Systematic Errors in Language Model Reasoning

    arXiv:2602.15868v3 Announce Type: replace Abstract: Large language models (LLMs) exhibit failure modes on seemingly trivial tasks. We propose a formalisation of LLM interaction using a deterministic multi-tape Turing machine, where each tape represents a distinct component: input…

  8. arXiv cs.AI TIER_1 English(EN) · Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran, Dung D. Le ·

    OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate para…

  9. arXiv cs.AI TIER_1 English(EN) · Cai Ke, Xinghao Chen, Xiaoyu Shen, Keyu Chen, Siyu An, Junnan Dong, Ruifeng Xu, Ruizhi Qiao, Xing Sun ·

    Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning

    arXiv:2609.18461v1 Announce Type: new Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence. While early flat retrieval methods score memory fragments independently and neglect …

  10. arXiv cs.AI TIER_1 English(EN) · Yizheng Yang, Haining Yu, Yuechen Wang, Yikai Hou, Xing Fu, Jinbo Yang, Tianqing Zhu ·

    First Token Matters: Understanding Safety Collapse in Large Reasoning Models

    arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preferen…

  11. arXiv cs.AI TIER_1 English(EN) · Rebecca Ansell, Autumn Toney-Wails ·

    Clueing up LLMs with Tool-Augmented Deductive Reasoning

    arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that require integrating evidence across multiple reasoning steps, ma…

  12. arXiv cs.AI TIER_1 English(EN) · Zhongdi Qu, Carla P. Gomes ·

    A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

    arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these observations with a mechanistic account. We show that the m…

  13. Hugging Face Daily Papers TIER_1 English(EN) ·

    When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

    Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading redu…

  14. arXiv cs.AI TIER_1 English(EN) · Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan ·

    Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation

    arXiv:2603.15547v2 Announce Type: replace-cross Abstract: Modeling student misconceptions in a realistic manner is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating distractor answers for multiple-c…

  15. arXiv cs.AI TIER_1 English(EN) · Zhiren Gong, Yikun Hou, Zihao Zeng, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim ·

    State of Thought Enables Endogenous Reasoning

    arXiv:2609.16055v1 Announce Type: cross Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reas…

  16. arXiv cs.AI TIER_1 English(EN) · Hung-Hsuan Chen ·

    Thinking Deeper, Not Longer: Memory-Efficient Test-Time Reasoning with Depth-Recurrent Transformers for Compositional Generalization

    arXiv:2603.21676v2 Announce Type: replace-cross Abstract: Standard Transformers have a fixed computational depth, limiting their ability to generalize to tasks that require variable-depth reasoning. The usual remedy, Chain-of-Thought (CoT), spends tokens to reason, inflating the …

  17. arXiv cs.AI TIER_1 English(EN) · Jinyang Zhang, Weibin Liao, Keqin Bao, Sihang Li, Shaobo Wang, Muyang Ye, Hongxin Ding, Yue Fang, Tianyi Tang, Fei Huang, Kexin Yang, Xingzhang Ren, Dayiheng Liu ·

    The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis

    arXiv:2609.16076v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge …

  18. arXiv cs.AI TIER_1 English(EN) · Pratibha Zunjare, Michael Hsiao ·

    NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect

    arXiv:2603.02504v3 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on natural language tasks but remain unreliable in mathematical reasoning, frequently generating fluent yet logically inconsistent solutions. We present \textbf{NeuroProlog…

  19. arXiv cs.AI TIER_1 English(EN) · Hongyu Gu, Chang Liu, Jingwen Fu ·

    How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition

    arXiv:2609.13747v1 Announce Type: new Abstract: Large language models solve hard problems through intermediate computations across multi-step reasoning. Traditional chain-of-thought encodes these computations as tokens. Recent continuous and recurrent methods instead move partial…

  20. arXiv cs.AI TIER_1 English(EN) · Simon Schug, Brenden M. Lake ·

    Thought without systematicity? Evaluating reasoning models on rule induction tasks

    arXiv:2609.13948v1 Announce Type: cross Abstract: A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so…

  21. arXiv cs.CL TIER_1 English(EN) · Fangan Dong, Zuming Yan, Xuri Ge, Zhiwei Xu, Mengqi Zhang, Xuanang Chen, Ben He, Xin Xin, Zhumin Chen, Ying Zhou ·

    Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering

    arXiv:2601.19847v3 Announce Type: replace Abstract: Despite the strong reasoning capabilities of recent large language models (LLMs), achieving reliable performance on challenging tasks often requires post-training or computationally expensive sampling strategies, limiting their …

  22. arXiv cs.CL TIER_1 English(EN) · Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar, Yannick Detrois, Mattia Nee, Pierre-Carl Langlais, Ivan P. Yamshchikov ·

    Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks

    arXiv:2504.07825v2 Announce Type: replace Abstract: Commonsense reasoning is a key language model capability, as it is purportedly a prerequisite for many basic tasks, unlike specific factual knowledge. It is often measured with multiple-choice questions (MCQ) benchmarks, e.g. He…

  23. arXiv cs.CL TIER_1 English(EN) · Emmy Liu, Graham Neubig, Jacob Andreas ·

    An Incomplete Loop: Deductive, Inductive, and Abductive Reasoning in Language Models

    arXiv:2404.03028v4 Announce Type: replace Abstract: Modern language models (LMs) can learn to perform new tasks in different ways: in instruction following, the target task is described explicitly in natural language; in few-shot prompting, the task is specified implicitly with a…

  24. arXiv cs.CL TIER_1 English(EN) · Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada, Sara Khosravi, Zeynep Senahan Yildiz, Amit Sahai, Nanyun Peng ·

    Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

    arXiv:2609.13773v1 Announce Type: cross Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human pre…

  25. arXiv cs.CL TIER_1 English(EN) · Runa Yoshida, Kosuke Nishida, Kyosuke Nishida ·

    Improving Mathematical Reasoning Capabilities in Large Language Models via Reasoning Process Error Classification

    arXiv:2609.15145v1 Announce Type: new Abstract: The reasoning ability of large language models (LLMs) is a critical factor for practical LLM-based applications. To investigate the current reasoning capability of LLMs, we clarify the types of errors that arise in LLMs' reasoning p…

  26. arXiv cs.AI TIER_1 English(EN) · Danchun Chen, Qiyao Yan, Chenpeng Wang, Liangming Pan ·

    Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models

    arXiv:2601.04260v2 Announce Type: replace Abstract: Understanding how Large Language Models (LLMs) perform logical reasoning internally remains a fundamental challenge. While prior mechanistic studies focus on identifying task specific circuits, they leave open the question of wh…

  27. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Adam Jatowt ·

    The Magnitude Mirage: Rethinking Confidence for Reasoning-Intensive Retrieval

    Many production RAG systems implement retrieval abstention by thresholding raw similarity scores, implicitly treating score magnitude as a confidence signal. We demonstrate that this practice degrades systematically as queries require reasoning beyond semantic matching. Across 11…

  28. arXiv cs.AI TIER_1 English(EN) · Yasmine Briefs, Christoph Weidenbach ·

    Extending SMT Solving with Non-Ground Clause Learning

    arXiv:2609.11509v1 Announce Type: new Abstract: Quantifier instantiation is currently the main approach to non-ground SMT solving: solvers generate ground instances and solve the resulting ground SMT problems with CDCL(T)-style reasoning. When a conflict is found, conflict analys…

  29. Hugging Face Daily Papers TIER_1 English(EN) ·

    Thought without systematicity? Evaluating reasoning models on rule induction tasks

    Reasoning models frequently fail on structurally equivalent task variants, indicating a lack of systematicity in their cognitive abilities.

  30. arXiv cs.AI TIER_1 English(EN) · Yiqi Li, Xu Chen, Chen Ju, Jiangchao Yao, Zhaoyang Li, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Yu Wang ·

    Structural Process Supervision for Latent Chain-of-Thought Reasoning

    arXiv:2609.09928v1 Announce Type: new Abstract: Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack direct process supervisi…

  31. arXiv cs.LG TIER_1 English(EN) · Deblina Kar ·

    A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

    arXiv:2609.10654v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a multi-stage rule-chaining framework that performs compositi…

  32. arXiv cs.CL TIER_1 English(EN) · Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng ·

    Negative Self-Distillation: Learning to Reason by Avoiding Flaws

    arXiv:2609.11699v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. …

  33. arXiv cs.AI TIER_1 English(EN) · Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao, Xiaoyuan Wang, Soroush Vosoughi ·

    Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

    arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already master…

  34. Hugging Face Daily Papers TIER_1 English(EN) ·

    Negative Self-Distillation: Learning to Reason by Avoiding Flaws

    On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can …

  35. arXiv cs.AI TIER_1 English(EN) · Kang Chen, Sihan Zhao, Yixin Cao, Yu-Gang Jiang ·

    From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts

    arXiv:2609.06403v1 Announce Type: new Abstract: Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effectiv…

  36. arXiv cs.AI TIER_1 English(EN) · Qihao Yuan ·

    Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths

    arXiv:2609.09001v1 Announce Type: new Abstract: Multi-step LLM reasoning lacks a machine-recheckable ledger: discarded reasoning paths leave no auditable record. We propose the Deposon scattering layer, which binds each node of an LLM-generated concept-decomposition graph to a tw…

  37. arXiv cs.AI TIER_1 English(EN) · Yu-Hang Wu, Yu-Jie Xiong, Henghua Zhang, Bairui Zhang, Jia-Chen Zhang, Shaohua Li ·

    Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models

    arXiv:2609.08186v1 Announce Type: new Abstract: The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed to enhance safety alignment, the stability of alignment mechanisms under extended…

  38. arXiv cs.AI TIER_1 English(EN) · Suhyeong Park, Junha Jung, Jaewoo Kang ·

    Reason Through the Latent! Making Latent Visual Reasoning Necessary

    arXiv:2609.06746v1 Announce Type: new Abstract: Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model …

  39. arXiv cs.LG TIER_1 English(EN) · Rafael Pardinas, Ehsan Kamalloo, David Vazquez, Alexandre Drouin ·

    Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning

    arXiv:2604.02007v3 Announce Type: replace Abstract: Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopted by frontier open-weight models. However, their training recipes and domain mixt…

  40. arXiv cs.LG TIER_1 Dansk(DA) · Baris Arat, Emre Sefer ·

    Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools

    arXiv:2602.18613v2 Announce Type: replace Abstract: Standard reranking evaluations study how a reranker orders candidates returned by an upstream retriever. This setup couples ranking behavior with retrieval quality, so differences in output cannot be attributed to the ranking po…

  41. arXiv cs.LG TIER_1 English(EN) · Jaewoo Lim, Sungbok Shin, Sanghyun Hong ·

    Do Reasoning Representations Help Humans Evaluate LLM Outputs?

    arXiv:2609.09038v1 Announce Type: new Abstract: Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether the…

  42. arXiv cs.LG TIER_1 English(EN) · Yuwen Hao, Menglin Yang ·

    Think Wider: Mitigating Latent Rank Collapse in Implicit Chain-of-Thought Reasoning

    arXiv:2609.07406v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rationales increase decoding length, latency, and context cost. Implicit CoT offers a mor…

  43. arXiv cs.LG TIER_1 English(EN) · Yimeng Ye, Shuang Chen, Wenxuan Huang, Manyuan Zhang, Kaituo Feng, Zhangquan Chen, Jiayu Chen, Yucheng Zhou, Yicheng Xiao, Zhiyuan Feng, Tianyu Shi ·

    Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

    arXiv:2609.07148v1 Announce Type: new Abstract: While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" …

  44. arXiv cs.CL TIER_1 English(EN) · Ryan Lail ·

    Decomposing LLM-Judge Uncertainty to Target Expert Labels

    arXiv:2609.06444v2 Announce Type: replace Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, a…

  45. arXiv cs.CL TIER_1 English(EN) · Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu ·

    ConvMem: Convolutional Memory for Long-Context Reasoning

    arXiv:2609.10441v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective …

  46. arXiv cs.AI TIER_1 English(EN) · Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang ·

    Boosting LLM Reasoning via Human-Inspired Reward Shaping

    arXiv:2602.04265v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, existing reward formulations typically treat exploration and conso…

  47. arXiv cs.AI TIER_1 English(EN) · Zuoyou Jiang, Li Zhao, Rui Sun, Ruohan Sun, Zhongjian Li, Jing Li, Daxin Jiang, Zuo Bai, Cheng Hua ·

    Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning

    arXiv:2512.23515v2 Announce Type: replace-cross Abstract: Signal decay and regime shifts pose recurring challenges for data-driven investment strategies in non-stationary markets, where conventional time-series and machine learning approaches often struggle to generalize beyond h…

  48. arXiv cs.AI TIER_1 Deutsch(DE) · Yufeng Zhao, Junnan Liu, Hongwei Liu, Dongsheng Zhu, Yuan Shen, Songyang Zhang, Kai Chen ·

    When Tools Hurt LLM Reasoning: State-Dependent Belief Revision under External Evidence

    arXiv:2508.15754v2 Announce Type: replace-cross Abstract: Tool use is often assumed to monotonically improve reasoning, where external evidence is expected to help when relevant and be ignored when irrelevant. We show that this assumption fails in a state-dependent way. Across be…

  49. arXiv cs.AI TIER_1 English(EN) · Andreas Reich, Claudia Thoms, Tobias Schrimpf ·

    Introducing HALC: A general pipeline for the systematic and reliable construction of prompts for automated coding with LLMs in the computational social sciences

    arXiv:2507.21831v2 Announce Type: replace-cross Abstract: LLMs are seeing widespread use for task automation, including automated coding in the social sciences. However, even though researchers have proposed different prompting strategies, their effectiveness varies across LLMs a…

  50. arXiv cs.AI TIER_1 English(EN) · Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder ·

    When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

    arXiv:2602.06286v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of unknown outcomes. However, it is unclear whether LLMs act as if they hold cohere…

  51. arXiv cs.AI TIER_1 English(EN) · Sining Zhoubian, Dan Zhang, Jie Tang ·

    ReST-RL: Reinforcing LLM Reasoning through Unified Self-Training and Value-Guided Search

    arXiv:2508.19576v3 Announce Type: replace Abstract: With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method - Group Relative Policy Optimization (GRPO) - has achieved critical success, yet it still suffers from the issue of …

  52. arXiv cs.AI TIER_1 English(EN) · Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He ·

    A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

    arXiv:2609.07821v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a princ…

  53. arXiv cs.AI TIER_1 English(EN) · Xiaodong Wang, Peixi Peng ·

    Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning

    arXiv:2609.07036v1 Announce Type: cross Abstract: We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations such as I'm doing, in contrast to the revision- and backtracking-oriented Aha Moment. We refer to their corresponding l…

  54. arXiv cs.AI TIER_1 English(EN) · Qi Wang, Chengcheng Wan, Jiangtao Wang ·

    SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure

    arXiv:2609.06540v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inferenc…

  55. arXiv cs.AI TIER_1 English(EN) · Mar Gonz\`alez I Catal\`a, Haitz S\'aez de Oc\'ariz Borde, Davide Murari, Carola-Bibiane Sch\"onlieb, Pietro Li\`o, George Monta\~nez ·

    Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

    arXiv:2609.09030v1 Announce Type: new Abstract: Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work …

  56. Hugging Face Daily Papers TIER_1 English(EN) ·

    Negative Self-Distillation: Learning to Reason by Avoiding Flaws

    Negative Self-Distillation improves large language model reasoning by pushing models away from self-generated flawed reasoning via a dynamic gating mechanism that protects linguistic capabilities.

  57. Hugging Face Daily Papers TIER_1 English(EN) ·

    ConvMem: Convolutional Memory for Long-Context Reasoning

    While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iterativel…

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

    A*-Thought-V2 models chain-of-thought reasoning as hidden-state trajectories to selectively retain explicit reasoning steps or compress them into continuous latent tokens, improving accuracy and efficiency.

  59. arXiv cs.CL TIER_1 English(EN) · Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu, Alice Oh, Taekyung Kim ·

    Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

    arXiv:2609.04753v1 Announce Type: new Abstract: Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about …

  60. arXiv cs.AI TIER_1 English(EN) · Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri ·

    Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

    arXiv:2609.05385v1 Announce Type: new Abstract: LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or dec…

  61. arXiv cs.AI TIER_1 English(EN) · Andrea Gregor de Varda, Sana Pandey, Pengrui Han, Jacob Andreas, Evelina Fedorenko ·

    Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning

    arXiv:2609.04463v1 Announce Type: cross Abstract: In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solve 'two plus five'. In contrast, LLMs are more brit…

  62. arXiv cs.AI TIER_1 English(EN) · Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane ·

    Extremely Sparse Supervision Incentivizes Reasoning Ability

    arXiv:2609.04565v1 Announce Type: new Abstract: Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning…

  63. arXiv cs.AI TIER_1 English(EN) · Weicai Huang (Beijing MQPat Technologies, Co., Ltd.) ·

    DODR: Deterministic Operator-Driven Reasoning in Latent Space

    arXiv:2609.04782v1 Announce Type: new Abstract: Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which induces three fundamental defects in complex logical reasoning: error accumulation, probability substituting necessity, and t…

  64. arXiv cs.AI TIER_1 English(EN) · Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou, Shao-Qun Zhang ·

    GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

    arXiv:2609.05284v1 Announce Type: new Abstract: Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs). However, the reasoning processes of LLMs often exhibit uncertainty, where LLMs often produce a proliferation of divergent branches …

  65. arXiv cs.LG TIER_1 English(EN) · Jeffrey Lai, Anthony Bao, John Quinn, William Gilpin ·

    Fractal basins trap latent reasoning

    arXiv:2609.04963v1 Announce Type: new Abstract: Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widel…

  66. arXiv cs.CL TIER_1 English(EN) · Shi-Qi Yan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling ·

    ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

    arXiv:2609.04648v1 Announce Type: new Abstract: Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong …

  67. Hugging Face Daily Papers TIER_1 English(EN) ·

    Revisiting Complete Reasoning Traces for Post-Training

    Large language models gain reasoning improvements from truncated trajectory endpoints rather than full reasoning traces, reducing redundancy while benefiting supervised fine-tuning and reinforcement learning.

  68. Hugging Face Daily Papers TIER_1 English(EN) ·

    Reason Through the Latent! Making Latent Visual Reasoning Necessary

    CVRR enforces recurrent hidden-state computation as the required image-conditioned pathway for visual reasoning, preserving model competence while distinguishing latent information from actual predictive use.

  69. arXiv cs.AI TIER_1 English(EN) · Yigit Utku Bulut ·

    It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

    arXiv:2609.03436v1 Announce Type: cross Abstract: Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both …

  70. arXiv cs.LG TIER_1 English(EN) · Yuxiang Wang, Kunyu Feng, Yingda Shen, Haoning Xu, Junyu Wang, Zhizheng Wu ·

    RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

    arXiv:2609.03379v1 Announce Type: new Abstract: Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two…

  71. arXiv cs.LG TIER_1 English(EN) · Leqi Zheng, Jinbo Su, Fang Niu, Chaokun Wang, Weiping Wang, Jiajun Zhang, Shannan Yan, Jie Wu, Zhaolu Kang, Rong Fu, Hang Zhang ·

    Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

    arXiv:2609.03342v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from …

  72. arXiv cs.LG TIER_1 English(EN) · Frank Hu, Shriram Chennakesavalu, David Graff ·

    Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

    arXiv:2609.03177v1 Announce Type: new Abstract: Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoni…

  73. arXiv cs.CL TIER_1 English(EN) · Harshad Khadilkar, Abhay Gupta ·

    Causal-Counterfactual RAG: The Integration of Causal-Counterfactual Reasoning into RAG

    arXiv:2509.14435v3 Announce Type: replace Abstract: Large language models (LLMs) have transformed natural language processing (NLP), enabling diverse applications by integrating large-scale pre-trained knowledge. However, their static knowledge limits dynamic reasoning over exter…

  74. arXiv cs.AI TIER_1 English(EN) · Hui Wu, Hengyi Cai, Jinman Zhao, Xinran Chen, Ziheng Li, Zhejun Zhao, Shuaiqiang Wang, Yuchen Li, Dawei Yin ·

    Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment

    arXiv:2602.01207v2 Announce Type: replace Abstract: Offline preference optimization aligns reasoning models from fixed chosen--rejected pairs, yet standard methods apply gradient updates from every pair regardless of its training value under the current policy. We argue that this…

  75. arXiv cs.AI TIER_1 English(EN) · Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Chu Liu, Pengjun Xie, Yilun Zhao, Shu Wu ·

    CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

    arXiv:2609.04083v1 Announce Type: cross Abstract: MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions w…

  76. arXiv cs.AI TIER_1 English(EN) · Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim ·

    </think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

    arXiv:2609.03633v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to …

  77. Hugging Face Daily Papers TIER_1 English(EN) ·

    Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

    Distinct reasoning operations in language models are geometrically separable in hidden representations, with structure emerging across layers and depending on contextual reasoning context.

  78. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Shu Wu ·

    CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

    MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating…

  79. Hugging Face Daily Papers TIER_1 English(EN) ·

    </think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

    Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that in…

  80. Hugging Face Daily Papers TIER_1 English(EN) ·

    Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

    Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, eit…

  81. arXiv cs.CL TIER_1 English(EN) · Xu Zou, Jie Tang ·

    Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

    arXiv:2609.02702v1 Announce Type: new Abstract: Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing t…

  82. arXiv cs.AI TIER_1 English(EN) · Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld, Joshua Tenenbaum, Kevin Ellis ·

    Induction and Inquiry via Probabilistic Reasoning over Language and Code

    arXiv:2609.01815v1 Announce Type: new Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) d…

  83. arXiv cs.AI TIER_1 English(EN) · Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal, Robert Kirk, John Hughes ·

    Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

    arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make…

  84. arXiv cs.AI TIER_1 English(EN) · Zhao Ji, Wenqing Chen, Zhixuan Chu, Jianxing Yu, Jingping Liu, Shanhe Zhao, Zibin Zheng ·

    SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning

    arXiv:2609.02336v1 Announce Type: new Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-bas…

  85. arXiv cs.AI TIER_1 English(EN) · Rongzhi Zhu, Yi Liu, Jiancheng Wang, Xiangyu Liu, Zequn Sun, Yiwei Wang, Yu Deng, Zijian Zhou, Wei Hu ·

    When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning

    arXiv:2505.15276v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to "overthink" leads to inefficiencies. Although "save-thinking" prompts are intended to mitigate this issue, we find that LRMs …

  86. Hugging Face Daily Papers TIER_1 English(EN) ·

    CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

    CORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.

  87. Hugging Face Daily Papers TIER_1 English(EN) ·

    Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

    A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to disti…

  88. arXiv cs.CL TIER_1 English(EN) · Zabir Al Nazi, Shubhashis Roy Dipta, Sudipta Kar ·

    DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems

    arXiv:2601.06853v3 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting is widely adopted for mathematical problem solving, including in low-resource languages, yet its behavior under irrelevant context remains underexplored. To systematically study this challenge, w…

  89. arXiv cs.LG TIER_1 English(EN) · Ngoc-Hieu Nguyen, Parshin Shojaee, Phuc Minh Nguyen, Nan Zhang, Chandan K Reddy, Khoa D Doan, Rui Zhang ·

    Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

    arXiv:2605.17026v2 Announce Type: replace Abstract: Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedures. While these methods reliably improve pass@1 …

  90. arXiv cs.LG TIER_1 English(EN) · Mariia Drozdova, Aidan Sirbu, Pietro Miotti, Robert Obryk, Mayalen Etcheverry, Eyvind Niklasson, Blake Richards ·

    Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

    arXiv:2609.01449v1 Announce Type: new Abstract: Diffusion models and recursive reasoners are both iterative, but they carry information across iterations differently. We add a persistent hidden state to a diffusion denoiser and remove its timestep conditioning, leaving a single s…

  91. arXiv cs.CL TIER_1 English(EN) · Jonathan Zheng, Zirui Shao, Alan Ritter, Wei Xu ·

    Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

    arXiv:2609.00184v1 Announce Type: new Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactu…

  92. arXiv cs.AI TIER_1 English(EN) · Jingcheng Yu, Mingliang Zeng, Qiwei Ye ·

    FloydNet: A Learning Paradigm for Global Relational Reasoning

    arXiv:2601.19094v3 Announce Type: replace-cross Abstract: Learning algorithmic computation often requires explicit relational intermediate states, yet many graph processors maintain their primary states on individual entities. We introduce \fnet and \textbf{Pivotal Attention} (PA…

  93. arXiv cs.AI TIER_1 English(EN) · Wenjing Zhang, Jiangze Yan, Jieyun Huang, Yi Shen, Shuming Shi, Ping Chen, Ning Wang, Zhaoxiang Liu, Kai Wang, Shiguo Lian ·

    HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation

    arXiv:2603.10359v2 Announce Type: replace Abstract: Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitations of rejection sampling. Standard methods treat the teacher as a static filter, discarding comple…

  94. arXiv cs.AI TIER_1 English(EN) · Ye Tian, Zihao Wang, Onat Gungor, Xiaoran Fan, Tajana Rosing ·

    LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning

    arXiv:2601.13880v2 Announce Type: replace Abstract: Personalized lifestyle health analysis requires long-horizon, multi-dimensional reasoning over heterogeneous lifestyle signals, and recent advances in mobile sensing and large language models (LLMs) make such support increasingl…

  95. arXiv cs.AI TIER_1 English(EN) · Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson ·

    HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning

    arXiv:2510.15144v4 Announce Type: replace Abstract: Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus…

  96. arXiv cs.AI TIER_1 English(EN) · Hamed Babaei Giglou, Jennifer D'Souza, S\"oren Auer ·

    Do General NLP Embeddings Capture Ontological Reasoning?

    arXiv:2609.00177v1 Announce Type: cross Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings disting…

  97. arXiv cs.AI TIER_1 English(EN) · Zhaoliang Chen, Jie Fu ·

    Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

    arXiv:2609.01117v1 Announce Type: new Abstract: Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - …

  98. arXiv cs.AI TIER_1 English(EN) · Lu Cheng ·

    Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs

    arXiv:2609.00738v1 Announce Type: new Abstract: Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored---a failure mode we call \textit{reasoning basin coll…

  99. arXiv cs.AI TIER_1 English(EN) · Beidi Zhao, Gexin Huang, Ciro Zhang, Anqi Li, Yusheng Tan, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li ·

    SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning

    arXiv:2609.00342v1 Announce Type: new Abstract: Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WS…

  100. Hugging Face Daily Papers TIER_1 English(EN) ·

    Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs

    Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored---a failure mode we call \textit{reasoning basin collapse}. We introduce BASIN, a training-free, stru…

  101. arXiv cs.LG TIER_1 English(EN) · Xin Jiang, Minhao Wang, Wen Wu, Zhentao Xie, Shangheng Du, Jinxin Shi, Jiabao Zhao ·

    ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

    arXiv:2608.28771v1 Announce Type: new Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong resul…

  102. arXiv cs.CL TIER_1 English(EN) · Zeming Chen, Angelika Romanou, Gail Weiss, Antoine Bosselut ·

    PERK: Long-Context Reasoning as Test-Time Learning

    arXiv:2507.06415v3 Announce Type: replace Abstract: Long-context reasoning requires accurately identifying relevant information in extensive, noisy input contexts. In this work, we propose PERK (Parameter Efficient Reasoning over Knowledge), a scalable approach for learning to en…

  103. arXiv cs.AI TIER_1 English(EN) · Swapnil Parekh, Naman Goyal ·

    Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning

    arXiv:2605.11467v2 Announce Type: replace-cross Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of *reasoning theater*: deliberative-looking steps that contribute nothing to correctness. This wastes inference tok…

  104. arXiv cs.AI TIER_1 English(EN) · Rubing Chen, Jian Wang, Wenjie Li, Xiao-Yong Wei, Qing Li ·

    To Retrieve or To Think? Cross-Boundary Context Evolution for Multi-hop Complex Reasoning

    arXiv:2601.08747v3 Announce Type: replace-cross Abstract: Current context augmentation methods, such as retrieval-augmented generation, play a crucial role in bridging a model's internal knowledge boundary and external evidence for multi-hop reasoning. However, they often follow …

  105. arXiv cs.AI TIER_1 English(EN) · Yadong Wang, Haodong Chen, Yu Tian, Chuanxing Geng, Dong Liang, Xiang Chen ·

    Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning

    arXiv:2602.01695v2 Announce Type: replace Abstract: Latent reasoning reduces the token-generation cost of chain-of-thought reasoning by replacing explicit intermediate tokens with continuous latent transitions. However, existing latent reasoning methods usually rely on dense and …

  106. arXiv cs.AI TIER_1 English(EN) · Qingjie Zhang, Yujia Fu, Yang Wang, Liu Yan, Tao Wei, Ke Xu, Minlie Huang, Han Qiu ·

    Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models

    arXiv:2509.24711v4 Announce Type: replace Abstract: Current answering paradigms for Large Reasoning Models (LRMs) often fail to account for the fact that some questions may lie beyond the model's operational capability boundary, leading to long but unproductive reasoning. In this…

  107. arXiv cs.AI TIER_1 English(EN) · Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi ·

    Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

    arXiv:2608.30256v1 Announce Type: cross Abstract: Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has oft…

  108. arXiv cs.AI TIER_1 English(EN) · Yangsong Lan, Renkai Hu, HongKai Zheng, Bo Zhang, Renzhi Wang, Hongliang Dai, Piji Li ·

    MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation

    arXiv:2608.29623v1 Announce Type: cross Abstract: Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains chall…

  109. arXiv cs.AI TIER_1 English(EN) · Dylan Jayabahu, Tinuade Adeleke ·

    The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

    arXiv:2608.28859v1 Announce Type: cross Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable v…

  110. arXiv cs.AI TIER_1 English(EN) · Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos ·

    Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

    arXiv:2608.31082v1 Announce Type: new Abstract: Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions fo…

  111. arXiv cs.LG TIER_1 English(EN) · Zimo Shi, Xander Tifft, Wen Xing ·

    Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steering

    arXiv:2608.29070v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. This assumes CoT surface…

  112. arXiv cs.AI TIER_1 English(EN) · Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo ·

    Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

    arXiv:2608.31075v1 Announce Type: new Abstract: Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extendi…

  113. arXiv cs.AI TIER_1 English(EN) · Jayanta Sadhu, Sayem Shahad, Kenneth Marino ·

    DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

    arXiv:2608.30413v1 Announce Type: new Abstract: Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in …

  114. arXiv cs.AI TIER_1 English(EN) · Hanlin Tian, Minhao Li, Yu Mi, Sihan Zhu, Zhao Yang, Yuxiang Wang, Hongquan Zhu, Qiufei Hu ·

    Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

    arXiv:2608.30322v1 Announce Type: new Abstract: Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that…

  115. arXiv cs.AI TIER_1 English(EN) · Yu Li, Wei Li, Xin Gao, Mengyuan Sun, Xiaoyang Wang, Qizhi Pei, Lijun Wu ·

    SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature

    arXiv:2608.30214v1 Announce Type: new Abstract: Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited e…

  116. arXiv cs.CL TIER_1 English(EN) · Mengdan Zhu, Senhao Cheng, Liang Zhao ·

    Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

    arXiv:2604.07518v2 Announce Type: replace Abstract: Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insuf…

  117. arXiv cs.CL TIER_1 English(EN) · Manas Mehta, Fangcong Yin, Greg Durrett ·

    Randomized YaRN Improves Length Generalization for Long-Context Reasoning

    arXiv:2606.23687v2 Announce Type: replace Abstract: Large language models (LLMs) are typically pretrained on short sequences and then extended to work on longer sequences with additional training. However, such LLMs still struggle to further generalize to very long sequences. We …

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

    Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

    Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in defeasible reasoning, the datasets have been sta…

  120. arXiv cs.LG TIER_1 English(EN) · ShengYun Peng, Pin-Yu Chen, Eric Smith, Song Jiang, Hongyuan Zhan, Haozhu Wang, Mahesh Pasupuleti, Duen Horng Chau, Jianfeng Chi ·

    Large Reasoning Models Learn Better Alignment from Flawed Thinking

    arXiv:2510.00938v3 Announce Type: replace Abstract: Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flaw…

  121. arXiv cs.CL TIER_1 English(EN) · Lucas Bandarkar, Alan Ansell, Trevor Cohn ·

    Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts

    arXiv:2603.17070v2 Announce Type: replace Abstract: In this work, we analyze shortcomings in cross-lingual knowledge transfer in large, modern reasoning LLMs. We demonstrate that the perceived gap in knowledge transfer is primarily a script barrier. First, we conduct an observati…

  122. Hugging Face Daily Papers TIER_1 English(EN) ·

    Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

    Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this tod…

  123. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

    This work proposes a structured ladder for scaling large reasoning models beyond human supervision by tracing autonomous rewards and self-generated experience, while identifying risks and evaluation dimensions.

  124. Hugging Face Daily Papers TIER_1 English(EN) ·

    Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

    A protocol separates task instructions from private convention artefacts to explicitly test agent dependence on hidden knowledge, validated by calibration tasks showing near-zero performance without access.

  125. Hugging Face Daily Papers TIER_1 English(EN) ·

    Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions

    Large language models encode whether structurally impossible math or code prompts are unanswerable via a hidden-state direction, but fail to abstain because this recognition signal is misaligned with safety-refusal pathways, indicating a routing rather than encoding failure.

  126. arXiv cs.AI TIER_1 English(EN) · Maciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler ·

    Performance Foundations of Parallel & Distributed Reasoning Language Models

    arXiv:2608.27046v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models (RL…

  127. arXiv cs.AI TIER_1 English(EN) · Sachin Gopal Wani, Ajay Dholakia, David Ellison ·

    The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

    arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric…

  128. arXiv cs.AI TIER_1 English(EN) · Haizhao Fan, Yuchi Xiong, Jize Wang, Xinping Guan, Xinyi Le ·

    SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers

    arXiv:2608.26836v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language understanding, yet they struggle with strict multi-step reasoning, frequently suffering from hallucinations and inconsistency. Existing soluti…

  129. arXiv cs.AI TIER_1 English(EN) · Zike Yuan, Han Zhang, Jianzhi Yan, Le Liu, Cai Ke, Huozhi Zhou, Jian Xie, Jiran Yin, Yukun Cao, Yue Yu, Hui Wang, Ming Liu, Bing Qin ·

    GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL

    arXiv:2608.27142v1 Announce Type: new Abstract: Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting…

  130. arXiv cs.AI TIER_1 English(EN) · Philipp Schr\"oppel ·

    Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning

    arXiv:2608.26166v1 Announce Type: cross Abstract: Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - intermediate steps toward solutions - open up high-stakes applications by enabling human inspect…

  131. arXiv cs.AI TIER_1 English(EN) · Yunxin Sun, Abulhair Saparov ·

    Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning

    arXiv:2509.03345v3 Announce Type: replace Abstract: Non-deductive reasoning, encompassing inductive and abductive reasoning, is essential in addressing complex real-world questions. One key feature of inductive and abductive reasoning is that there are many valid hypotheses; the …

  132. arXiv cs.AI TIER_1 English(EN) · Yuxiang Wang, Junhao Gan, Shengxiang Gao, Shenghao Ye, Zhengyi Yang, Jianzhong Qi ·

    Beyond Linearization: Attributed Table Graphs for Table Reasoning

    arXiv:2601.08444v2 Announce Type: replace Abstract: Table reasoning, a task to answer questions by reasoning over data presented in tables, is an important topic due to the prevalence of knowledge stored in tabular formats. Recent solutions use Large Language Models (LLMs) for th…

  133. arXiv cs.AI TIER_1 English(EN) · Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He ·

    From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning

    arXiv:2505.22203v3 Announce Type: replace-cross Abstract: Trustworthy verifiers are essential for the success of reinforcement learning with verifiable reward (RLVR), which is the core methodology behind various large reasoning models such as DeepSeek-R1. In complex domains like …

  134. arXiv cs.CL TIER_1 English(EN) · Yuxin Zi, Cong Xu, Suparna Bhattacharya, Martin Foltin, Amit Sheth ·

    Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification

    arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed, mathe…

  135. arXiv cs.CL TIER_1 English(EN) · Fei Ding ·

    The Thousand-Graph Hypothesis: A Testable Hypothesis of Task-Conditioned Relation Materialization in Repository-Level Code Reasoning

    arXiv:2608.26602v1 Announce Type: cross Abstract: Large software repositories are often beyond model context limits. Training repository knowledge into models is costly and quickly stale, while local retrieval can miss scattered requirements, and explicit relation graphs add ongo…

  136. arXiv cs.LG TIER_1 English(EN) · Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang ·

    Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

    arXiv:2608.27351v1 Announce Type: new Abstract: Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to …

  137. Hugging Face Daily Papers TIER_1 English(EN) ·

    Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

    Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group …

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL

    Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting topological structures from noisy text is highl…

  139. Hugging Face Daily Papers TIER_1 English(EN) ·

    Performance Foundations of Parallel & Distributed Reasoning Language Models

    Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models (RLMs) such as DeepSeek-R1, o3, and Kimi k1.5 show th…

  140. Hugging Face Daily Papers TIER_1 English(EN) ·

    SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers

    Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language understanding, yet they struggle with strict multi-step reasoning, frequently suffering from hallucinations and inconsistency. Existing solutions like Chain-of-Thought (CoT) lack rigorous ve…

  141. arXiv cs.LG TIER_1 English(EN) · Pin-Han Ho, Limei Peng, Yiming Miao, Yan Jiao ·

    Epistemic Memory: A Validity Layer for Self-Maintaining Intelligent Systems

    arXiv:2510.16899v2 Announce Type: replace Abstract: AI memory mechanisms primarily focus on preserving information content, often neglecting the validity conditions under which knowledge remains applicable, leading to semantic coordinate drift when agents move, change sensors, or…

  142. arXiv cs.LG TIER_1 English(EN) · Samuel Lippl, Thomas McGee, Kimberly Lopez, Ziwen Pan, Pierce Zhang, Salma Ziadi, Oliver Eberle, Ida Momennejad ·

    AlgoTrace: Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models

    arXiv:2510.15987v3 Announce Type: replace Abstract: How do inference time and latent computations enable large language models (LLMs) to solve multi-step reasoning problems? We introduce AlgoTrace, a framework for tracing and steering algorithmic operations in the model latent sp…

  143. arXiv cs.CL TIER_1 English(EN) · Chung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei Weng ·

    ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability

    arXiv:2510.09062v2 Announce Type: replace Abstract: Recent advances in long chain-of-thought (CoT) reasoning have largely prioritized answer accuracy and token efficiency, while overlooking aspects critical to trustworthiness. We argue that usable reasoning systems must be trustw…

  144. arXiv cs.CL TIER_1 English(EN) · Srimonti Dutta, Akshata Kishore Moharir ·

    Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

    arXiv:2608.26036v1 Announce Type: cross Abstract: Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability c…

  145. arXiv cs.CL TIER_1 English(EN) · Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen, Yu Yang ·

    Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

    arXiv:2608.25542v1 Announce Type: cross Abstract: Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection …

  146. arXiv cs.CL TIER_1 English(EN) · Rongchen Zhao, Yu Chen, Juyuan Wang, Zhouting Mo, Jianxing Yu, Wenqing Chen, Jingping Liu ·

    PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning

    arXiv:2608.25486v1 Announce Type: cross Abstract: Long Narrative Reasoning is an essential capability for processing and reasoning over complex narratives. While retrieval-augmented generation provides a promising framework, existing methods still face two critical challenges: co…

  147. arXiv cs.CL TIER_1 English(EN) · Lam So, Canhui Wu, Han Lin ·

    GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning

    arXiv:2608.25583v1 Announce Type: new Abstract: Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference cost and latency. In contrast, instruction-tuned mod…

  148. arXiv cs.CL TIER_1 English(EN) · Jinpu Jiang, Xuan Wu, Wenhao Song, Bo Yang, You Zhou, Hongwei Ge, Heow Pueh Lee, Yanchun Liang, Chunguo Wu ·

    ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains

    arXiv:2608.25487v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in…

  149. arXiv cs.CL TIER_1 English(EN) · Casey Kennington ·

    A Primer on Computational Semantics for Artificial Intelligence Systems

    arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed abo…

  150. Hugging Face Daily Papers TIER_1 English(EN) ·

    Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

    Evolution strategies improve reasoning diversity and Pass@K over GRPO through sparse functional updates and population diversity, supporting a hybrid training approach.

  151. Hugging Face Daily Papers TIER_1 English(EN) ·

    Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

    Code-as-World represents physical environments as executable code to enable quantitative reasoning and scalable supervision for vision-language models.

  152. Hugging Face Daily Papers TIER_1 English(EN) ·

    Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

    Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation re…

  153. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Chunguo Wu ·

    ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains

    Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in news and social media poses a serious challenge…

  154. arXiv cs.AI TIER_1 English(EN) · Zhen Bi, Xueshu Chen, Yan Wang, Zhizhi Peng, Haosen Hong, Zhen Wang, Zhixuan Chu, Bingyu Zhu, Jungang Lou ·

    Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

    arXiv:2608.23982v1 Announce Type: new Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. Conditional memory provides an explicit lookup pathway that complements dense neural representa…

  155. arXiv cs.AI TIER_1 English(EN) · Zhenyu Wu, Siyuan Chen, Changchun Yang, Jiaqi Dong, Min Zhou, Ali Almadan, Talal Hammad, Faisal Wahbo, Aminullah Tora, Mona Alshahrani, Xin Gao ·

    TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

    arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchm…

  156. arXiv cs.AI TIER_1 English(EN) · Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han, Ligeng Zhu ·

    Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

    arXiv:2608.24658v1 Announce Type: new Abstract: Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for …

  157. arXiv cs.AI TIER_1 English(EN) · Shaojie Zhu, Zhaobin Wang, Chengxiang Zhuo, Hui Lu, Bo Hu, Zang Li ·

    Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs

    arXiv:2312.17535v2 Announce Type: replace Abstract: In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking has led to large language models (LLMs) attracting widespread attention. However, restricted by expensive costs, many studies have to…

  158. arXiv cs.LG TIER_1 English(EN) · Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong, Fumiya Uchiyama, Kenji Kubo, Kohei Hayashi, Masahiro Suzuki, Yutaka Matsuo ·

    Steering Recurrent Reasoners at Inference Time with Readout Feedback

    arXiv:2608.24136v1 Announce Type: new Abstract: Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more st…

  159. arXiv cs.AI TIER_1 English(EN) · Md Mahadi Hasan Nahid, Davood Rafiei ·

    PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding

    arXiv:2608.24082v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for r…

  160. arXiv cs.AI TIER_1 English(EN) · Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie ·

    Recursive Agentic Reasoning

    arXiv:2608.23956v1 Announce Type: new Abstract: Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are often evaluated in isolation, making their gains difficult to compare across models, benchmarks, and evaluation pipelines. We introdu…

  161. Hugging Face Daily Papers TIER_1 English(EN) ·

    PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding

    Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for reasoning. Existing approaches typically reason ove…

  162. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Davood Rafiei ·

    PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding

    Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for reasoning. Existing approaches typically reason ove…

  163. arXiv cs.CL TIER_1 English(EN) · Murat Dura, Serkan \"Ozt\"urk, Selma Tekir ·

    Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

    arXiv:2608.22332v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we i…

  164. arXiv cs.AI TIER_1 English(EN) · Yujie Zhang, Bin Gao, Tulika Mitra ·

    SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning

    arXiv:2608.21614v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity t…

  165. arXiv cs.AI TIER_1 English(EN) · Orion Powers, Daniella Seum, Khaled Slhoub ·

    More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

    arXiv:2608.22048v1 Announce Type: new Abstract: Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy while fewer quantify local runtime and energy, characterize failure modes, or ap…

  166. arXiv cs.AI TIER_1 English(EN) · Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman, Afsaneh Fatemi ·

    Small Reasoning Models are Instruction Followers in Function Calling

    arXiv:2608.22472v1 Announce Type: new Abstract: Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs function-calling accuracy through fine-tuning, reinforcement learning (RL), and multi-agent fra…

  167. arXiv cs.AI TIER_1 English(EN) · Min Chen, Shengjun Zhang, Yuxin Li, Zhang Zhang, Xin Fei, Chong Xia, Yueqi Duan ·

    ParallelWorld: Test-Time Scaling for Embodied Reasoning

    arXiv:2608.22971v1 Announce Type: new Abstract: Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of e…

  168. arXiv cs.AI TIER_1 English(EN) · Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu, Ziyi Wang, Xun Zhao, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen ·

    Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

    arXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method train…

  169. arXiv cs.AI TIER_1 English(EN) · Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang, Di Wang ·

    Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

    arXiv:2608.23497v1 Announce Type: new Abstract: Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious c…

  170. arXiv cs.AI TIER_1 English(EN) · Weihang Pan, Zhengxu Yu, Yuxiang Zhang, Wenzhi Li, Zhongming Jin, Binbin Lin, Xiaofei He, Jieping Ye ·

    ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

    arXiv:2608.21860v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoning. However, advanced Large Reasoning Models (LRMs…

  171. arXiv cs.AI TIER_1 English(EN) · Lihui Liu, Zihao Wang, Hanghang Tong ·

    Neural-Symbolic Reasoning over Knowledge Graphs: A Survey from a Query Perspective

    arXiv:2412.10390v2 Announce Type: replace Abstract: Knowledge graph reasoning is pivotal in various domains such as data mining, artificial intelligence, the Web, and social sciences. These knowledge graphs function as comprehensive repositories of human knowledge, facilitating t…

  172. arXiv cs.AI TIER_1 English(EN) · Zhejian Lai, Xiang Geng, Zhijun Wang, Yang Bai, Jiahuan Li, Rongxiang Weng, Jingang Wang, Xuezhi Cao, Xunliang Cai, Shujian Huang ·

    AdaR: A Framework for Equipping LLMs with Adaptive Reasoning

    arXiv:2510.04617v3 Announce Type: replace Abstract: Mathematical reasoning is a primary indicator of large language models (LLMs) intelligence. However, existing LLMs exhibit failures in robustness and generalization. This paper attributes these deficiencies to spurious reasoning…

  173. arXiv cs.CL TIER_1 English(EN) · Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu ·

    Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

    arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a large…

  174. arXiv cs.CL TIER_1 English(EN) · Yi Lu, Deyang Kong, Jianing Wang, Linsen Guo, Xue Wang, Qi Guo, Tao Gui, Xuanjing Huang, Wei Ye, Shikun Zhang, Wei Wang ·

    Adaptive Test-Time Compute Allocation for Block Diffusion Language Models in Complex Reasoning

    arXiv:2602.09555v3 Announce Type: replace Abstract: Recent advances in block diffusion language models have demonstrated competitive performance and strong scalability on reasoning tasks. However, their test-time compute allocation remains largely unexplored, leaving a critical s…

  175. Hugging Face Daily Papers TIER_1 English(EN) ·

    Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

    Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safety of LLM reasoning. Cross-a…

  176. Hugging Face Daily Papers TIER_1 English(EN) ·

    ParallelWorld: Test-Time Scaling for Embodied Reasoning

    Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of embodied reasoning from static perception toward …

  177. arXiv cs.CL TIER_1 English(EN) · Ravisri Valluri, Tung Nguyen, Aditya Grover ·

    Self-Speculation for Faster Reasoning Models

    arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor …

  178. arXiv cs.AI TIER_1 English(EN) · Haorui Xu, Yuzhou Zhu, Liyuan Gao ·

    DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

    arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steerin…

  179. arXiv cs.AI TIER_1 English(EN) · Said Slaoui ·

    S-AI-Recursive: Convergent Recursive Reasoning

    arXiv:2605.13872v2 Announce Type: replace-cross Abstract: This article introduces S-AI-Recursive, a bio-inspired Sparse Artificial Intelligence architecture in which reasoning is implemented as a hormonally regulated closed-loop iteration rather than a single feed-forward pass. T…

  180. arXiv cs.CL TIER_1 English(EN) · Lan Zhang, Marco Valentino, Jordan Meadows, Andre Freitas ·

    Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

    arXiv:2506.10903v2 Announce Type: replace Abstract: Statement autoformalization plays a crucial role in formal mathematical reasoning by enabling the automatic translation of natural language statements into formal languages. While recent advances using large language models (LLM…

  181. arXiv cs.CL TIER_1 English(EN) · Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang, Tingwen Liu ·

    Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

    arXiv:2608.21265v1 Announce Type: new Abstract: Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may …

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

    Mixed supervised fine-tuning on combined reasoning corpora outperforms next-chunk reinforcement learning in efficiency and final accuracy across mathematical and out-of-domain tasks.

  183. arXiv cs.CL TIER_1 English(EN) · Philipp Hellwig, Willem Zuidema, Claire E. Stevenson, Martha Lewis ·

    Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning

    arXiv:2604.06501v2 Announce Type: replace-cross Abstract: Analogical reasoning is a hallmark of human intelligence, enabling us to solve new problems by transferring knowledge from one situation to another. Yet, developing artificial intelligence systems capable of robust human-l…

  184. arXiv cs.LG TIER_1 English(EN) · Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi ·

    Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    arXiv:2608.19669v1 Announce Type: cross Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these…

  185. arXiv cs.AI TIER_1 English(EN) · Muhan Gao, Zih-Ching Chen, Kuan-Hao Huang ·

    The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning

    arXiv:2605.10828v2 Announce Type: replace Abstract: As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understanding how distracting information affects long-context performance becomes criti…

  186. arXiv cs.AI TIER_1 English(EN) · Gijs Kassenaar, Zhao Yang, Vincent Fran\c{c}ois-Lavet ·

    Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

    arXiv:2608.20256v1 Announce Type: new Abstract: Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on d…

  187. arXiv cs.AI TIER_1 English(EN) · Yiting Qu, Ziqing Yang, Chi Cui, Ye Leng, Junjie Chu, Yang Zhang ·

    EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

    arXiv:2608.20055v1 Announce Type: cross Abstract: Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be directly extracted from black-box models remains large…

  188. arXiv cs.AI TIER_1 English(EN) · Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao ·

    Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents

    arXiv:2608.19794v1 Announce Type: new Abstract: The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI). This paper reviews the evolution of LLM-center…

  189. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward fee…

  190. arXiv cs.AI TIER_1 English(EN) · Pratik Ghawate ·

    FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

    arXiv:2608.18534v1 Announce Type: new Abstract: Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagno…

  191. arXiv cs.LG TIER_1 English(EN) · Lirui Luo, Guoxi Zhang, Hongming Xu, Rongqing Li, Cong Fang, Lifeng Fan ·

    Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

    arXiv:2608.18574v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study contin…

  192. arXiv cs.CL TIER_1 English(EN) · Yajie Yin ·

    Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning

    arXiv:2608.19009v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literatur…

  193. arXiv cs.AI TIER_1 English(EN) · Gilad Abiri, Emanuel V. Towfigh ·

    Epistemic Subordination: Generative AI and the Infrastructure of Knowledge

    arXiv:2608.18758v1 Announce Type: cross Abstract: Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full bread…

  194. arXiv cs.AI TIER_1 English(EN) · Zuocheng Ying, Yang Yang, Yumou Wu, Chuanbo Zhu, Jiarui Wang, Ziqi Wu, Jingming Cai, Junqing Yu, Zikai Song ·

    From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

    arXiv:2608.18581v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. Existing end-to-end methods entangle …

  195. arXiv cs.AI TIER_1 English(EN) · Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang ·

    Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

    arXiv:2608.18419v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with w…

  196. arXiv cs.AI TIER_1 English(EN) · Zijie Meng, Xiwei Dai, Yixuan Tang, Jin Hao, Yang Feng, Fudong Zhu, Xiaoqiang Liu, Shaosheng Cao, Zuozhu Liu ·

    DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning

    arXiv:2608.18878v1 Announce Type: new Abstract: Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D den…

  197. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zuozhu Liu ·

    DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning

    Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D dental data. Most existing dental AI systems remain…

  198. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Pratik Ghawate ·

    FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

    Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase ord…

  199. Hugging Face Daily Papers TIER_1 English(EN) ·

    FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

    Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase ord…

  200. arXiv cs.CL TIER_1 English(EN) · Eduardo S\'anchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer ·

    The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

    arXiv:2608.18011v1 Announce Type: new Abstract: Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-A…

  201. arXiv cs.AI TIER_1 English(EN) · Yeabin Moon ·

    The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

    arXiv:2608.16956v1 Announce Type: new Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reason…

  202. arXiv cs.AI TIER_1 English(EN) · Xing Wei, Changmeng Zheng, XiaoYong Wei, Xiufen Ye, Qing Li ·

    DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation

    arXiv:2608.17282v1 Announce Type: new Abstract: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentrali…

  203. arXiv cs.AI TIER_1 English(EN) · Guozheng Sun ·

    SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

    arXiv:2608.17301v1 Announce Type: new Abstract: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their applica…

  204. arXiv cs.AI TIER_1 English(EN) · Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li ·

    Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

    arXiv:2608.17253v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward).…

  205. Hugging Face Daily Papers TIER_1 English(EN) ·

    Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

    Co-RL enables unsupervised reasoning via cooperative multi-agent reinforcement learning with peer-derived rewards, improving performance across text and vision tasks without ground-truth labels.

  206. arXiv cs.AI TIER_1 English(EN) · Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu ·

    Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

    arXiv:2604.06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and…

  207. arXiv cs.CL TIER_1 English(EN) · Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li ·

    $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

    arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not…

  208. arXiv cs.LG TIER_1 English(EN) · Jules Soria, Alban Grastien, Romain Xu-Darme, Julien Girard-Satabin, Zakaria Chihani, Daniela Cancila ·

    Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

    arXiv:2608.16773v1 Announce Type: new Abstract: Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic…

  209. arXiv cs.CL TIER_1 English(EN) · Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li ·

    Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

    arXiv:2608.16554v1 Announce Type: new Abstract: Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the…

  210. arXiv cs.AI TIER_1 English(EN) · Lirui Teng ·

    GRIP: Grounded Reasoning via Information-Restricted Premises

    arXiv:2608.16776v1 Announce Type: new Abstract: High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence functionally irrelevant. We call this failure mode query dominance. To address it, we introduce \…

  211. arXiv cs.AI TIER_1 Español(ES) · Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen ·

    ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

    arXiv:2608.16425v1 Announce Type: new Abstract: Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these para…

  212. arXiv cs.AI TIER_1 English(EN) · Bo Wen, Yuhao Chen, Erhan Bilal, Carla Agurto Rios, Chen Wang, Junchen Jiang ·

    Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis

    arXiv:2608.15303v1 Announce Type: new Abstract: Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase pri…

  213. arXiv cs.AI TIER_1 English(EN) · Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang ·

    The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

    arXiv:2608.14558v1 Announce Type: new Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative p…

  214. arXiv cs.AI TIER_1 English(EN) · Shufeng Kong, Xiaochuan Zhang, Caihua Liu ·

    Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration

    arXiv:2608.14569v1 Announce Type: new Abstract: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the mode…

  215. Hugging Face Daily Papers TIER_1 English(EN) ·

    SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

    Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relat…

  216. Hugging Face Daily Papers TIER_1 Español(ES) ·

    ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

    ParaTempo improves parallel reasoning efficiency by using temporal confidence to dynamically prune, retire, and reallocate reasoning branches without synchronization.

  217. arXiv cs.AI TIER_1 English(EN) · Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn, Soheil Feizi ·

    Early Stopping for Large Reasoning Models via Confidence Dynamics

    arXiv:2604.04930v2 Announce Type: replace-cross Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key chal…

  218. arXiv cs.AI TIER_1 English(EN) · Feng Xiong, Leyan Xue, Hongyu Lin ·

    Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners

    arXiv:2607.28336v3 Announce Type: replace Abstract: On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subsequent reasoning. Perception Success Rate (PSR), e…

  219. arXiv cs.AI TIER_1 English(EN) · Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen ·

    The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

    arXiv:2509.14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.…

  220. arXiv cs.AI TIER_1 English(EN) · Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig ·

    Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

    arXiv:2608.13760v1 Announce Type: cross Abstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces loo…

  221. arXiv cs.AI TIER_1 English(EN) · Cuong Dang, Hoang Anh Just, Ruoxi Jia ·

    Capacity-Dependent Effects of Data Selection for Reasoning

    arXiv:2608.13721v1 Announce Type: cross Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that …

  222. arXiv cs.AI TIER_1 English(EN) · Dayuan Zhao, Shengcao Cao, Yu-Xiong Wang, Liang-Yan Gui ·

    Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

    arXiv:2608.13570v1 Announce Type: cross Abstract: Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing rea…

  223. arXiv cs.AI TIER_1 English(EN) · Zhelun (Allen), Wu ·

    Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

    arXiv:2608.13926v1 Announce Type: new Abstract: Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent …

  224. arXiv cs.AI TIER_1 English(EN) · Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz ·

    BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs

    arXiv:2608.14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and mitigating bias, it largely focuses on final outputs…

  225. arXiv cs.AI TIER_1 English(EN) · Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi, David Lo ·

    Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

    arXiv:2608.13667v1 Announce Type: new Abstract: LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is confined to the Thought phase: while the agent serializes an action and waits for the environment, its reasoning is fro…

  226. arXiv cs.AI TIER_1 English(EN) · Panjing He, Mingyue Cheng, Yucong Luo, Li Li, Xiaohan Zhang ·

    SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning

    arXiv:2608.14452v1 Announce Type: new Abstract: Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet reasoning remains challenging for large language models (LLMs). Real-world workbooks often contain implicit cross-tabl…

  227. Hugging Face Daily Papers TIER_1 English(EN) ·

    $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

    In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same mo…

  228. Hugging Face Daily Papers TIER_1 English(EN) ·

    R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

    R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks.

  229. arXiv cs.AI TIER_1 English(EN) · Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin ·

    Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities

    arXiv:2608.12374v1 Announce Type: cross Abstract: Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation…

  230. arXiv cs.AI TIER_1 English(EN) · Rachel Lawrence, Jacqueline Maasch ·

    Position: Reasoning is a Learnable Rule-Based Process

    arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite im…

  231. arXiv cs.AI TIER_1 English(EN) · Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar ·

    Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

    arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning beh…

  232. arXiv cs.AI TIER_1 English(EN) · Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini ·

    Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

    arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-pro…

  233. arXiv cs.AI TIER_1 English(EN) · Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang ·

    ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

    arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent …

  234. arXiv cs.AI TIER_1 English(EN) · Yubo Li, Ramayya Krishnan, Rema Padman ·

    LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

    arXiv:2608.12321v1 Announce Type: cross Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditiona…

  235. arXiv cs.CL TIER_1 English(EN) · Gaurav Najpande, Tampu Ravi Kumar, Manan Roy Choudhury, Neha Valeti, Yanjie Fu, Vivek Gupta ·

    QUIETT: Query-Independent Table Transformation for Robust Reasoning

    arXiv:2602.20017v2 Announce Type: replace Abstract: Real-world tables often contain schema inconsistencies, heterogeneous value formats, and implicit relational structures that degrade table reasoning and question answering. Existing approaches address these issues at query time,…

  236. arXiv cs.AI TIER_1 English(EN) · Obed Junias, Maria Leonor Pacheco ·

    From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

    arXiv:2608.12836v1 Announce Type: cross Abstract: Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEI…

  237. Hugging Face Daily Papers TIER_1 English(EN) ·

    Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

    Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indistinguishable at the point of …

  238. Hugging Face Daily Papers TIER_1 English(EN) ·

    Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

    Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy.

  239. arXiv cs.AI TIER_1 English(EN) · Atahan Dokme, Benjamin Reichman, Larry Heck ·

    TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

    arXiv:2604.07801v2 Announce Type: replace-cross Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does emot…

  240. arXiv cs.AI TIER_1 English(EN) · Valentin Rodionov, Shamil Assylbekov ·

    TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

    arXiv:2608.11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literatur…

  241. arXiv cs.AI TIER_1 English(EN) · Suyash Mishra ·

    Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning

    arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that…

  242. arXiv cs.AI TIER_1 English(EN) · Yoshinori Watanabe ·

    The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

    arXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data di…

  243. Hugging Face Daily Papers TIER_1 English(EN) ·

    From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

    A framework decomposes compound logical options into atomic judgments and uses constrained optimization to improve reasoning over AND, OR, and NEITHER/NOR operators.

  244. Hugging Face Daily Papers TIER_1 English(EN) ·

    Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

    Reasoning training amplifies deliberative behaviors like self-correction more than high-correctness behaviors such as confidence calibration, revealing a gap between amplified and correctness-linked reasoning patterns.

  245. arXiv cs.AI TIER_1 English(EN) · Rose Niousha, Minwoo Kang, Narges Norouzi ·

    INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

    arXiv:2608.10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as ev…

  246. arXiv cs.AI TIER_1 English(EN) · Si'an Xie (Beijing University of Posts and Telecommunications), Jiaxun Liu (Peking University), Biao Yang (Kuaishou Technology), Wei Yuan (Kuaishou Technology), Fan Yang (Kuaishou Technology), Tingting Gao (Kuaishou Technology), Ming Wu (Beijing Universi… ·

    From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

    arXiv:2608.10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unex…

  247. arXiv cs.AI TIER_1 English(EN) · Tughanbulut Kurtulush ·

    When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

    arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds o…

  248. arXiv cs.AI TIER_1 English(EN) · Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat, Soumyabrata Pal, Ramasuri Narayanam, Dinesh Manocha ·

    ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling

    arXiv:2608.10928v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminis…

  249. arXiv cs.AI TIER_1 English(EN) · Xin Xu ·

    Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects

    arXiv:2608.10420v1 Announce Type: new Abstract: Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of valu…

  250. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Shamil Assylbekov ·

    TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

    Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literature, a capability that has not yet been directly mea…

  251. arXiv cs.AI TIER_1 English(EN) · Minhan Cho, Jimin Kweon ·

    Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing

    arXiv:2608.08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models…

  252. arXiv cs.LG TIER_1 English(EN) · Zachary Horvitz, Raghav Singhal, Hao Zou, Carles Domingo-Enrich, Zhou Yu, Rajesh Ranganath, Kathleen McKeown ·

    Rethinking Reasoning with MDLMs: Early Exits, Post-hoc Reasoning, and Beyond

    arXiv:2510.19990v2 Announce Type: replace Abstract: The reasoning paradigm, where language models reason before answering, has enabled breakthroughs on tasks such as mathematical problem-solving. While current tooling for reasoning is built around next-token prediction trained mo…

  253. arXiv cs.CL TIER_1 English(EN) · Hexuan Wang, Yaxuan Ren, Srikar Bommireddypalli, Shuxian Chen, Adarsh Prabhudesai, Rongkun Zhou, Elina Baral, Philipp Koehn ·

    SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation

    arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To enable fine-grained diagnostic analysis beyond end-task accuracy, SciTaRC pairs ea…

  254. arXiv cs.CL TIER_1 English(EN) · Tianjun Zhong, Linyang He, Ziyang Li, Nima Mesgarani ·

    From Chains to DAGs: Probing the Graph Structure of Reasoning in LLMs

    arXiv:2601.17593v3 Announce Type: replace Abstract: Recent progress in large language models has renewed interest in how multi-step reasoning is represented internally. While prior work often treats reasoning as a linear chain, many reasoning problems can be more naturally modele…

  255. arXiv cs.CL TIER_1 English(EN) · Mengxi Xiao, Kailai Yang, Pengde Zhao, Enze Zhang, Ziyan Kuang, Zhiwei Liu, Weiguang Han, Shu Liao, Lianting Huang, Guojun Xiong, Victor Gutierrez Basulto, Jinpeng Hu, Min Peng, Qianqian Xie, Sophia Ananiadou ·

    MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy

    arXiv:2512.09636v3 Announce Type: replace Abstract: Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into interpretations, decisions, or claims whose spec…

  256. arXiv cs.AI TIER_1 English(EN) · Hongyi James Cai, Junlin Wang, Xiaoyin Chen, Bhuwan Dhingra ·

    How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

    arXiv:2505.24273v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yielding significant improvements on challenging reasoning tasks through extended chains of…

  257. arXiv cs.AI TIER_1 English(EN) · Muhammad Ali Shafique, Kelly Marchisio ·

    Hidden Language Consistency Phenomena in Reasoning LLMs

    arXiv:2608.08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingua…

  258. arXiv cs.AI TIER_1 English(EN) · Zhengze Huang, Luyang Yu, Di Hong, Xinzhe Huang, Wanyu Lin, Zhixuan Chu, Zhan Qin, Tianhang Zheng ·

    REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

    arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where fl…

  259. arXiv cs.AI TIER_1 English(EN) · Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi ·

    Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions

    arXiv:2608.07968v1 Announce Type: cross Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency cons…

  260. arXiv cs.AI TIER_1 English(EN) · Juncheng Dong, Ding Tong, Ishan Gupta, Yuyan Wang ·

    LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing

    arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality…

  261. Hugging Face Daily Papers TIER_1 English(EN) ·

    Thought-Level Beam Search for Reasoning

    Gambit improves reasoning model efficiency by using thought-level beam search to dynamically allocate compute to promising reasoning traces under fixed hardware budgets.

  262. Hugging Face Daily Papers TIER_1 English(EN) ·

    StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

    Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, thereby discarding the supervisi…

  263. LessWrong (AI tag) TIER_1 English(EN) · star2vec ·

    No sign of backtracking in latent reasoning: the final answer simply settles in instead

    <img alt="fig0_banner_dense.png" src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1789247969/lexical_client_uploads/nbvufluvrxky0i10mb5d.png" /><p><span style="white-space: pre-wrap;">Solving a hard math problem is not linear, it's trial and error.</span></p><p><span s…

  264. LessWrong (AI tag) TIER_1 English(EN) · Chi Nguyen ·

    Perspectives in favour of improving conceptual reasoning capabilities

    <p><span style="white-space: pre-wrap;">Below are some informal perspectives on improving AIs’ conceptual reasoning capabilities that motivate work in the area. Conceptual reasoning here means reasoning about questions where we cannot verify the answer, also don’t have data on un…

  265. arXiv cs.CV TIER_1 English(EN) · Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu ·

    Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

    arXiv:2608.27549v1 Announce Type: new Abstract: Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit represent…

  266. LessWrong (AI tag) TIER_1 English(EN) · Shahriar Golchin ·

    Do AI Models Want to Be Monitored? Measuring Monitorability Disposition in Large Reasoning Models

    <p><span>As AI models take on increasingly high-stakes responsibilities, understanding what a model is actually doing instead of simply constraining its outputs has become one of the most important safety challenges in AI.&nbsp;</span><a href="https://arxiv.org/abs/2512.18311?ref…

  267. arXiv cs.CV TIER_1 English(EN) · Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap ·

    Investigating Relational Reasoning in VLMs

    arXiv:2608.23518v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we…

  268. arXiv cs.CV TIER_1 English(EN) · Zesheng Yang, Lingling Zhang, Xinyu Zhang, Cheng Zhang, Pengyu Li, Heng Wang, Lin Wu ·

    GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning

    arXiv:2608.15517v1 Announce Type: new Abstract: Chain-of-thought reasoning has substantially improved the problem-solving capabilities of multimodal large language models. Fine-grained visual evidence, however, remains difficult to preserve and reuse across text-based reasoning s…

  269. arXiv cs.CV TIER_1 English(EN) · Xiaohan Zhang, Feng Gu, Xudong Rao, Xuhao Pan, Tao Wei, Zhou Pan, Kun Zhan ·

    ChainSpace: A Chained-Reasoning Paradigm for Spatial Intelligence

    arXiv:2608.15788v1 Announce Type: new Abstract: Spatial intelligence requires foundation models to maintain coherent spatial state across interactions with the physical world. However, existing data-centric approaches typically treat spatial reasoning as independent question-answ…

  270. LessWrong (AI tag) TIER_1 English(EN) · Chi Nguyen ·

    Introducing the Conceptual Reasoning Index

    <p><span>Associated announcement tweet.</span></p><p><span>We are planning to release blog posts properly arguing the case for this kind of work in the future.</span></p><h1><span>tl;dr</span></h1><p><span>A core hope for managing AI risks is that AIs will help us understand the …

  271. HN — anthropic stories TIER_1 English(EN) · optimalsolver ·

    Anthropic: Introducing The Conceptual Reasoning Index

  272. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    Create a Reasoning-Focused LLM: A Practical Guide to Streaming, Curating, and Fine-Tuning the SupraLabs Reasoning Corpus

    <p>This tutorial provides a complete workflow for building a compact, reasoning-focused language model. By streaming the SupraLabs reasoning corpus from Hugging Face, we apply quality filters and curate data for Supervised Fine-Tuning (SFT). Using SmolLM2-135M-Instruct and LoRA, …

  273. Medium — fine-tuning tag TIER_1 English(EN) · Sowmithdurusoju ·

    Creating Reasoning Models with GRPO Finetuning

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sowmithdurusoju/creating-reasoning-models-with-grpo-finetuning-20a36b0e57a9?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1200/1*52togsagBUurrab71b6N4A.gif" widt…

  274. Medium — fine-tuning tag TIER_1 English(EN) · Omokemiayodimeji ·

    Hop-1: Distilling One Reasoning Step Into a 270-Million-Parameter Model

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@omokemiayodimeji/hop-1-distilling-one-reasoning-step-into-a-270-million-parameter-model-189aa724e545?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1600/1*t7YE56T…

  275. Towards AI TIER_1 English(EN) · Vedant Pandhare ·

    RLVR: From Human Feedback to Verifiable Truth: How Modern LLMs Actually Learn to Reason

    <p><em>14 min read · AI Research · LLM Training</em></p><p>I want to tell you about the moment I stopped thinking about AI as a prediction machine.</p><p>It happened when I was reading through the DeepSeek-R1 paper at an ungodly hour — the kind of reading you do when something ge…

  276. Medium — fine-tuning tag TIER_1 English(EN) · Okan Yenigün ·

    Fine-Tuning Qwen3–4B with Teacher-Generated Reasoning

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/fine-tuning-qwen3-4b-with-teacher-generated-reasoning-01d48c05e852?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2600/1*sAOLPgpTmbU8oKoRH9cplw.jpeg" width…

  277. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    First Token Matters: Understanding Safety Collapse in Large Reasoning Models

    <h1> First Token Matters: Understanding Safety Collapse in Large Reasoning Models </h1> <p>The recent shift toward Large Reasoning Models (LRMs) — models that explicitly "think" through a Chain-of-Thought (CoT) before providing an answer — has promised a new era of complex proble…

  278. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    The LLM Knowledge-Reasoning Tradeoff: Why 2026's Best Models Are Deliberately Fact-Minimized — And Faster Than Ever

    <h2> Table of Contents </h2> <ol> <li>The Model That Spent 21 Minutes Drawing a Circle</li> <li>The LLM Knowledge-Reasoning Tradeoff: The Physics Behind It</li> <li>Why Reasoning Compresses Better Than Facts</li> <li>The Hallucination Paradox: World-Class Reasoning, High Factual …

  279. dev.to — LLM tag TIER_1 English(EN) · mayankpallai ·

    How LLMs Learned to Reason: SFT --> RLHF --> RLVR

    <h2> 1. The Starting Line: The Last Non-Reasoning Flagships </h2> <p>GPT-4.5, DeepSeek-V3, and Claude 3.5 Sonnet share something that has nothing to do with benchmark scores: they were the last major models built entirely on the "pretrain, then instruct-tune" recipe. All internal…

  280. dev.to — LLM tag TIER_1 English(EN) · Patience Sitati ·

    When RAG Fails: Anchoring Bias and Metacognition in Modern LLMs

    <p><em>## How a bizarre Claude Code CLI bug exposed the limits of multi-turn troubleshooting loops, and why the future of AI relies on cognitive routing.</em></p> <p>I recently ran into an issue with Claude Code on Windows.</p> <p>I gave Claude everything it needed: the terminal …

  281. r/MachineLearning TIER_1 English(EN) · /u/Typical-Scene-5794 ·

    Latent Reasoning Landscape in 2026: Mapping BDH-CQ, HRM/TRM, Coconut [D]

    <!-- SC_OFF --><div class="md"><p>After following various arXiv papers and researcher discussions on X/bluesky about latent reasoning and continual learning, one idea which resonates strongly is that path forward (towards AGI) may depend less on generating ever-longer chains of t…

  282. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Prime Agent: Scaling Long-Horizon Reasoning via Recursive Subagents and Persistent REPLs

    <h1> Prime Agent: Scaling Long-Horizon Reasoning via Recursive Subagents and Persistent REPLs </h1> <p>The evolution of agentic AI has reached a critical bottleneck. While frontier models like GPT-4 and Claude 3.5 Sonnet have demonstrated impressive reasoning capabilities, their …

  283. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Next-Chunk Reasoning: Why RL Might Not Actually Beat SFT for no-CoT Data

    <h1> Next-Chunk Reasoning: Why RL Might Not Actually Beat SFT for no-CoT Data </h1> <p>The transition from standard Supervised Fine-Tuning (SFT) to Reinforcement Learning (RL) has been hailed as the primary catalyst for the reasoning capabilities observed in recent frontier model…

  284. r/LocalLLaMA TIER_1 English(EN) · /u/devildip ·

    Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vvc6pw/qwen_35_4b_iq2_xs_1667_reasoning_performance_from/"> <img alt="Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation" src="https://preview.redd.it/69geg9n38xkh1.png?width=140&a…

  285. r/MachineLearning TIER_1 English(EN) · /u/aglet_factorial ·

    Epistemic Intelligence in Machine Learning Neurips Workshop page limit? [D]

    <!-- SC_OFF --><div class="md"><p>I'm aiming to submit a paper to The 3rd Workshop on Epistemic Intelligence in Machine Learning at Neurips<br /> <a href="https://eiml.cc/">https://eiml.cc/</a></p> <p>I can't find a page limit anywhere on their website and I've emailed the organi…

  286. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    TR-GRPO Stabilizes Reasoning Training by Regulating Token Gradients

    <h1> TR-GRPO Stabilizes Reasoning Training by Regulating Token Gradients </h1> <p><em>Billing Support — August 21, 2026</em></p> <h2> The failure: unlikely tokens can control the update </h2> <p>Group Relative Policy Optimization, or GRPO, is a critic-free reinforcement-learning …

  287. r/LocalLLaMA TIER_1 English(EN) · /u/SandyL925 ·

    SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vu5dzi/sensenova_u15lite_full_release_expert_training/"> <img alt="SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference" src="https://preview.redd.it/lenqskwdhnkh1.png?w…

  288. dev.to — LLM tag TIER_1 English(EN) · Arisyn ·

    Beyond SQL Accuracy: Building Evidence Chains for AI Data Agents

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj6adfu5xhoaud2tgmnnl.png"><img alt=" " height="512" …

  289. dev.to — LLM tag TIER_1 English(EN) · Ken W Alger ·

    The Reasoning Ledger: Remembering Decisions, Not Just Data

    <p><em>Part 4 of the Building the AI Memory Stack series</em></p> <p>After finishing the previous article, I looked at the repository a little differently. The specifications were still there. The Architecture Decision Records were still there. The glossary entries were still the…

  290. r/MachineLearning TIER_1 English(EN) · /u/moschles ·

    BDH-CQ: IN-CONTEXT LEARNING WITH RECURRENT LATENT REASONING [R]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1vov5r5/bdhcq_incontext_learning_with_recurrent_latent/"> <img alt="BDH-CQ: IN-CONTEXT LEARNING WITH RECURRENT LATENT REASONING [R]" src="https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnCYTqbJkElIQ…

  291. r/Anthropic TIER_1 English(EN) · /u/Advanced-Cat9927 ·

    Beyond Prompt Engineering: Building an Auditable Epistemic State Machine

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1w4s6t4/beyond_prompt_engineering_building_an_auditable/"> <img alt="Beyond Prompt Engineering: Building an Auditable Epistemic State Machine" src="https://external-preview.redd.it/4lHjeGaPeUUlkP4OWfiT2s4yOpK3k…

  292. Mastodon — mastodon.social TIER_1 Русский(RU) · [email protected] ·

    It is shown that intermediate tokens (ITG) in "reasoning models" cannot be reliably interpreted as "traces of thoughts". Although training with intermediate tokens improves

    Показывается, что промежуточные токены (ITG) в «моделях рассуждений» нельзя надежно трактовать как «следы мыслей». Хотя обучение с промежуточными токенами повышает точность, семантической связи между содержимым трасс и реальными логическими вычислениями модели часто нет. В задача…

  293. Mastodon — mastodon.social TIER_1 English(EN) · lucashendren ·

    Nice reality check for the 'just add more reasoning' reflex. On the IOL linguistic reasoning challenge, chain-of-thought made models perform worse, and constrai

    Nice reality check for the 'just add more reasoning' reflex. On the IOL linguistic reasoning challenge, chain-of-thought made models perform worse, and constraining output to JSON was the single most damaging choice. The failure mode isn't lack of steps, it's that formatting and …

  294. r/OpenAI TIER_2 English(EN) · /u/Advanced-Cat9927 ·

    Beyond Prompt Engineering: Building an Auditable Epistemic State Machine

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1w4s7a0/beyond_prompt_engineering_building_an_auditable/"> <img alt="Beyond Prompt Engineering: Building an Auditable Epistemic State Machine" src="https://external-preview.redd.it/4lHjeGaPeUUlkP4OWfiT2s4yOpK3kbu9…

  295. r/ClaudeAI TIER_2 English(EN) · /u/Enough-Plantain2785 ·

    Anthropic: Introducing The Conceptual Reasoning Index

    <!-- SC_OFF --><div class="md"><p><a href="https://alignment.anthropic.com/2026/conceptual-reasoning-index/">https://alignment.anthropic.com/2026/conceptual-reasoning-index/</a></p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://www.reddit.com/user/Enough-Plantain…

  296. r/singularity TIER_2 English(EN) · /u/EducationalCicada ·

    Anthropic: Introducing The Conceptual Reasoning Index

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/EducationalCicada"> /u/EducationalCicada </a> <br /> <span><a href="https://alignment.anthropic.com/2026/conceptual-reasoning-index/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/singularity/comments…