PulseAugur
EN
LIVE 10:24:19

New research tackles LLM reasoning, long-context, and tool integration

Multiple research papers explore advancements in large language model (LLM) reasoning capabilities, focusing on improving performance in long-horizon tasks and tool integration. Apple's research introduces LEAD, a method to overcome a "no-recovery bottleneck" in LLM reasoning by incorporating future validation and overlapping rollouts, enabling models to solve complex problems like Checkers Jumping. Other papers present TurnSight for turn-level hindsight self-distillation in tool-integrated reasoning, methods for measuring explainer stability, and techniques for analyzing and improving reasoning processes through test-time scaling and reflective learning from failed trajectories. Additionally, research investigates optimizing reasoning interfaces for shorter response times and evaluating long-context reasoning with novel memory mechanisms. AI

IMPACT These advancements aim to improve LLM reliability, efficiency, and capability in complex reasoning tasks, potentially enabling more sophisticated AI applications.

RANK_REASON Multiple academic papers published on arXiv detailing new methods and analyses for LLM reasoning.

Read on Apple Machine Learning Research →

AI-generated summary · Google Gemini · from 228 sources. How we write summaries →

New research tackles LLM reasoning, long-context, and tool integration

COVERAGE [228]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning

    Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposition is essential for stability, extreme decomposition creates a “no-recovery bot…

  2. arXiv cs.CL TIER_1 English(EN) · Bingxuan Li, Simo Du, Yue Guo ·

    Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent

    arXiv:2604.07269v2 Announce Type: replace Abstract: Clinical expertise improves not only by acquiring medical knowledge, but by accumulating experience that yields reusable diagnostic patterns. Recent LLMs-based diagnostic agents have shown promising progress in clinical reasonin…

  3. arXiv cs.CL TIER_1 English(EN) · Hamed Damirchi, Ignacio Meza De la Jara, Damith Ranasinghe, Yuhang Liu, Javen Shi ·

    Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

    arXiv:2608.05660v1 Announce Type: cross Abstract: As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. Recent trajectory-based methods seek th…

  4. arXiv cs.CL TIER_1 English(EN) · Xinye Wang, Junxiao Liu, Shujian Huang ·

    RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

    arXiv:2608.06347v1 Announce Type: new Abstract: Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising paradigm…

  5. arXiv cs.CL TIER_1 English(EN) · Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han ·

    On-Policy Delta Distillation for Multilingual Math Reasoning

    arXiv:2608.05802v1 Announce Type: new Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Pol…

  6. arXiv cs.CL TIER_1 English(EN) · Hongbo Ma, Bangji Yang, Yunqian Selina Cheng, Jiajun Fan, Hanwen Zhang, Ge Liu ·

    Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving

    arXiv:2608.05254v1 Announce Type: new Abstract: Large language models can derive a plausible mathematical object yet still violate explicit requirements--for example, by omitting a modular reduction, returning a non-integer, or using the wrong encoded answer form. We introduce Co…

  7. arXiv cs.CL TIER_1 English(EN) · Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs ·

    Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

    arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles t…

  8. arXiv cs.AI TIER_1 English(EN) · Emanuele Marconato, Samuele Bortolotti, Emile van Krieken, Paolo Morettin, Elena Umili, Antonio Vergari, Efthymia Tsamoura, Andrea Passerini, Stefano Teso ·

    Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

    arXiv:2510.14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or structural constraints. As such, it represents one of the most promising avenues for reliable a…

  9. arXiv cs.AI TIER_1 English(EN) · James Xu Zhao, Bryan Hooi, See-Kiong Ng ·

    Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet

    arXiv:2509.06861v3 Announce Type: replace Abstract: Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. However, frontier models still suffer from factuality hallucinations, raising the…

  10. arXiv cs.AI TIER_1 English(EN) · Hao Ai ·

    Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models

    arXiv:2608.05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, w…

  11. arXiv cs.AI TIER_1 English(EN) · ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng ·

    DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

    arXiv:2608.06243v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-…

  12. arXiv cs.AI TIER_1 English(EN) · Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen ·

    Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

    arXiv:2608.05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning dive…

  13. arXiv cs.AI TIER_1 English(EN) · Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Yang Li, Deguo Xia, Jizhou Huang ·

    Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

    arXiv:2608.05168v1 Announce Type: new Abstract: Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise from localized reasoning bugs in intermediate steps rather than from global incompetence. We…

  14. arXiv cs.AI TIER_1 English(EN) · Yu Gu, Zhi Zheng, Yunpeng Ba, Xialiang Tong, Mingxuan Yuan, Zhenkun Wang ·

    Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

    arXiv:2608.05541v1 Announce Type: new Abstract: Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such h…

  15. arXiv cs.LG TIER_1 English(EN) · Zonghuan Xu ·

    Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction

    arXiv:2608.06262v1 Announce Type: new Abstract: Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We f…

  16. arXiv cs.AI TIER_1 English(EN) · Byoungjae Min, Kennedy Edemacu, Sae-Hong Cho, Yoonhyuk Choi, Beakcheol Jang, Jong Wook Kim ·

    When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

    arXiv:2608.04591v1 Announce Type: cross Abstract: Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, th…

  17. arXiv cs.AI TIER_1 English(EN) · Qiyuan Zhu, Dezhi Li, Pengyu Cheng, Tianle Chen, Jiacheng Wang, Ruijie Shen, Hao Gu, Sida Lin, Zirui Liu, Jiacheng Liu, Sirui Han ·

    Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning

    arXiv:2608.04771v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inference cost. KV-cache compression is a common soluti…

  18. arXiv cs.CL TIER_1 English(EN) · Damien Sileo, Valentin Lacombe, Dimitri Kachler ·

    Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

    arXiv:2608.05148v1 Announce Type: new Abstract: Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tuning. We introduce Reasoning Core, a collection of 50 generators spanning mathema…

  19. arXiv cs.CL TIER_1 English(EN) · Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora ·

    Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

    arXiv:2608.05139v1 Announce Type: new Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill…

  20. arXiv cs.CL TIER_1 English(EN) · Ian B. de Haan, Peter van der Putten, Max van Duijn ·

    Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

    arXiv:2608.04646v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via …

  21. arXiv cs.CL TIER_1 English(EN) · Bhiman Kumar Baghel, Anna Chrabaszcz, Tessa Warren, Michael Walsh Dickey, Haley C. Dresang, Xiang Lorraine Li ·

    STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

    arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled eve…

  22. arXiv cs.AI TIER_1 English(EN) · Junqi Chen, Sirui Chen, Chaochao Lu ·

    Can Post-Training Transform LLMs into Causal Reasoners?

    arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in this domain, their precise causal estimation capabilities are still limited, and…

  23. arXiv cs.AI TIER_1 English(EN) · Purbesh Mitra, Sennur Ulukus ·

    Chained Recursive Language Models for Multi-Iteration Reasoning

    arXiv:2608.05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the fin…

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    ChronoVision: Temporal Reasoning via Latent State Reconstruction

    Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulat…

  25. Hugging Face Daily Papers TIER_1 English(EN) ·

    On-Policy Delta Distillation for Multilingual Math Reasoning

    On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical…

  26. Hugging Face Daily Papers TIER_1 English(EN) ·

    Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

    Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose step…

  27. arXiv cs.AI TIER_1 English(EN) · Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary ·

    Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

    arXiv:2608.04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single traje…

  28. arXiv cs.LG TIER_1 English(EN) · Jian Zhang, Bingyi Wang, Yizhi Liu ·

    CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning

    arXiv:2608.03673v1 Announce Type: new Abstract: Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning. Al…

  29. arXiv cs.CL TIER_1 English(EN) · Ming Shen, Zhikun Xu, Jacob Dineen, Xiao Ye, Ben Zhou ·

    BOW: Training Language Models to Reason Over Plausible Next Words

    arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word reasoning methods make this tension explicit: the…

  30. arXiv cs.AI TIER_1 English(EN) · Daeyoung Roh, Donghee Han ·

    Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG

    arXiv:2608.02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a proced…

  31. arXiv cs.AI TIER_1 English(EN) · Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang, Chen Zhang, Shunian Chen, Caijun Yan, Jianyao Xu, Shunyu Liu, Weijie Fu, Peiliang Li, Xiaozhi Chen, Yuxiang Cai ·

    SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

    arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mism…

  32. arXiv cs.AI TIER_1 English(EN) · Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu ·

    TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

    arXiv:2608.04007v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit ass…

  33. arXiv cs.AI TIER_1 English(EN) · Francesca Carlon, Vincent Ginis, Andres Algaba ·

    Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

    arXiv:2608.03401v1 Announce Type: cross Abstract: Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluat…

  34. arXiv cs.AI TIER_1 English(EN) · Shashwat Sourav, Aishwarya Balwani ·

    The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

    arXiv:2608.03291v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning…

  35. arXiv cs.AI TIER_1 English(EN) · Simon Lam-Muir ·

    The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner

    arXiv:2608.03263v1 Announce Type: cross Abstract: We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal training data. We grow an independent realization of a published 30M-parameter …

  36. arXiv cs.AI TIER_1 English(EN) · Dawei Liu, Haixu Song, Shuang Cheng, Shijie Wang, Haozheng Hou, Kaifeng Liu, Ermo Hua, Zhonghang Yuan, Zhijie Zhong, Yuchen Fan, Biqing Qi, Bowen Zhou ·

    PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory

    arXiv:2608.03048v1 Announce Type: cross Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irr…

  37. arXiv cs.AI TIER_1 English(EN) · Eddie Conti, \'Alvaro Parafita, Axel Brando ·

    Measuring Explainer Stability via Attribution Separability

    arXiv:2608.02697v1 Announce Type: cross Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definitio…

  38. arXiv cs.AI TIER_1 English(EN) · Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua ·

    ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

    arXiv:2608.03972v1 Announce Type: new Abstract: On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the exper…

  39. Hugging Face Daily Papers TIER_1 English(EN) ·

    Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

    Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose step…

  40. Hugging Face Daily Papers TIER_1 English(EN) ·

    ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

    On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-…

  41. Hugging Face Daily Papers TIER_1 English(EN) ·

    CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning

    Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning. Although large language models perform well on suc…

  42. Hugging Face Daily Papers TIER_1 English(EN) ·

    When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

    Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a…

  43. arXiv cs.CL TIER_1 English(EN) · Sean Gip Lim, William Chandra Tjhi, Hai Leong Chieu ·

    Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

    arXiv:2608.00533v1 Announce Type: new Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex lo…

  44. arXiv cs.CL TIER_1 English(EN) · Kang Liu, Zijing Wang, Yongkang Liu, Mengjie Zhao, Xiaocui Yang, Shi Feng, Yifei Zhang, Daling Wang ·

    TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

    arXiv:2608.01922v1 Announce Type: new Abstract: Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as reasoning trajectories grow, models may become less effective at using informat…

  45. arXiv cs.LG TIER_1 English(EN) · Hector Zenil, Luan Ozelim ·

    Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard

    arXiv:2608.01575v1 Announce Type: new Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning …

  46. arXiv cs.CL TIER_1 English(EN) · Zhaoxin Yu, Qi Shen, Hengli Li, Zhaowei Zhang, Song-Chun Zhu, Chi Zhang, Zilong Zheng ·

    GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

    arXiv:2608.02585v1 Announce Type: cross Abstract: Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these sta…

  47. arXiv cs.CL TIER_1 English(EN) · Mengting Ai, Jingrui He, Yue Guo ·

    Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression

    arXiv:2608.01631v1 Announce Type: new Abstract: KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that supports it. We test this assumption for large reasoning models and show that it c…

  48. arXiv cs.CL TIER_1 English(EN) · Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao ·

    Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

    arXiv:2608.02442v1 Announce Type: cross Abstract: Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify …

  49. arXiv cs.CL TIER_1 English(EN) · Shigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan, Anbang Yao ·

    Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization

    arXiv:2608.01078v1 Announce Type: new Abstract: We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to exhibit chain-of-thought reaso…

  50. Hugging Face Daily Papers TIER_1 English(EN) ·

    TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

    Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy s…

  51. Hugging Face Daily Papers TIER_1 English(EN) ·

    ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

    On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-…

  52. Hugging Face Daily Papers TIER_1 English(EN) ·

    When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

    Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a…

  53. Latent Space (swyx) TIER_1 English(EN) ·

    The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.

  54. arXiv cs.AI TIER_1 English(EN) · Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin ·

    On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

    arXiv:2607.29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, model…

  55. arXiv cs.CL TIER_1 English(EN) · Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang, Kun Liang, Weijie Liu, Yunfang Wu ·

    Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

    arXiv:2607.29185v1 Announce Type: new Abstract: Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilistic reward modeling, they rely on superficial cues t…

  56. arXiv cs.AI TIER_1 English(EN) · Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng ·

    ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

    arXiv:2607.28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not…

  57. arXiv cs.AI TIER_1 English(EN) · Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley ·

    How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

    arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into…

  58. arXiv cs.CL TIER_1 English(EN) · Sara Candussio, Daniel Scalena, Luca Bortolussi, Elisabetta Fersini, Malvina Nissim, Gabriele Sarti ·

    Demystifying Entropy-based Selection for Chain-of-Thought Compression in Large Reasoning Models

    arXiv:2607.28707v1 Announce Type: new Abstract: Entropy-based pruning has been proposed as an effective method for compressing Chain-of-Thought (CoT) reasoning with negligible accuracy loss. We test the robustness of low- and high-entropy CoT step selection methods across various…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard

    Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning benchmark that supplies one exactly. Using the T…

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

    Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded to…

  61. arXiv cs.CL TIER_1 English(EN) · Yecheng Wu, Song Han, Han Cai ·

    Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

    arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the demonstrati…

  62. arXiv cs.LG TIER_1 English(EN) · Songshuo Lu, Zhi Chen, Yaohua Tang ·

    Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

    arXiv:2607.27770v1 Announce Type: new Abstract: A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore b…

  63. arXiv cs.CL TIER_1 English(EN) · Haojin Wang, Yike Wang, Shangbin Feng, Hannaneh Hajishirzi, Yulia Tsvetkov ·

    MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning

    arXiv:2602.05307v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve strong performance by producing long chains of thought, but their inference costs are high and often generate redundant reasoning. Small language models (SLMs) are far more efficient, yet st…

  64. arXiv cs.CL TIER_1 English(EN) · Zheng Wu, Chenhao Xue, Shijie Zheng, Yijie Lu, Cheng Yang, Zhuosheng Zhang ·

    Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

    arXiv:2607.28478v1 Announce Type: new Abstract: As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a …

  65. arXiv cs.CL TIER_1 English(EN) · Amruta Parulekar, Jinu Lee, Dilek Hakkani-T\"ur, Hari Sundaram ·

    Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

    arXiv:2607.27783v1 Announce Type: new Abstract: Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously …

  66. arXiv cs.AI TIER_1 English(EN) · Vivek Shukla, Varun Shukla, Atul, Divya Mishra, Mehul Kumar Das ·

    A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

    arXiv:2607.26102v1 Announce Type: cross Abstract: Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally …

  67. arXiv cs.CL TIER_1 English(EN) · Arnav Hiray, Agam Shah, Caleb Lu, Meghaj Tarte, Harsit Mittal, Sudheer Chava ·

    Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

    arXiv:2607.26952v1 Announce Type: new Abstract: We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers nat…

  68. arXiv cs.CL TIER_1 English(EN) · Xuyao Feng, Anthony Hunter ·

    Making Implicit Premises Explicit in Logical Understanding of Enthymemes

    arXiv:2603.06114v2 Announce Type: replace Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language processing (NLP) methods for handling enthymemes can potentially identify enthymemes i…

  69. arXiv cs.CL TIER_1 English(EN) · Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, Aishwarya Balwani ·

    Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

    arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage…

  70. Hugging Face Daily Papers TIER_1 English(EN) ·

    Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

    As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bi…

  71. Hugging Face Daily Papers TIER_1 English(EN) ·

    Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

    We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. W…

  72. arXiv cs.AI TIER_1 English(EN) · Jianshuo Dong, Yujia Fu, Chuanrui Hu, Chao Zhang, Han Qiu ·

    Towards Understanding the Cognitive Habits of Large Reasoning Models

    arXiv:2506.21571v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monitoring model behaviors. Inspired by the obse…

  73. arXiv cs.CL TIER_1 English(EN) · Zhichao Yan, Yunxiao Zhao, Jiapu Wang, Jiaoyan Chen, Xiaoli Li, Ru Li, Jeff Z. Pan ·

    Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore

    arXiv:2601.15050v5 Announce Type: replace Abstract: Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from \textit{factual myopia}: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation. This driv…

  74. Hugging Face Daily Papers TIER_1 English(EN) ·

    Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

    Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language models can serve as an actionable signa…

  75. arXiv cs.AI TIER_1 English(EN) · Ahmed Haj Ahmed, Alvin Grissom II ·

    ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

    arXiv:2607.23058v1 Announce Type: cross Abstract: Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-…

  76. arXiv cs.AI TIER_1 English(EN) · Zirong Chen, Meiyi Ma ·

    Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

    arXiv:2607.23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic …

  77. arXiv cs.LG TIER_1 English(EN) · Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto ·

    Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

    arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on…

  78. arXiv cs.CL TIER_1 English(EN) · Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin ·

    LeAct: Learning to Reason from Expert Actions

    arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classic…

  79. arXiv cs.CL TIER_1 English(EN) · Xilun Chen, Ilia Kulikov, Vincent-Pierre Berges, Barlas O\u{g}uz, Rulin Shao, Gargi Ghosh, Jason Weston, Wen-tau Yih ·

    Learning to Reason for Factuality

    arXiv:2508.05618v2 Announce Type: replace Abstract: Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than their non-reasoning counterparts on long-form fac…

  80. Hugging Face Daily Papers TIER_1 English(EN) ·

    Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning

    Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four repres…

  81. arXiv cs.AI TIER_1 English(EN) · Anthony Zhan ·

    Simple Policy Gradients for Reasoning with Diffusion Language Models

    arXiv:2510.04019v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-training techniques, including reinforcement learning (RL), remains a key challenge for …

  82. arXiv cs.AI TIER_1 English(EN) · Kyle Cox, Darius Kianersi, Adri\`a Garriga-Alonso ·

    Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers

    arXiv:2603.01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisio…

  83. arXiv cs.AI TIER_1 English(EN) · Manqing Liu, David Williams-King, Ida Caspary, Linh Le, Hannes Whittingham, Puria Radmard, Cameron Tice, Edward James Young ·

    Diagnosing Pathological Chain-of-Thought in Reasoning Models

    arXiv:2602.13904v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT reasoning may exhibit failure modes that we note as pathologies, which prevent …

  84. arXiv cs.AI TIER_1 English(EN) · Renuka Oladri, Niveda Jawahar, Abdirisak Mohamed ·

    Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

    arXiv:2607.21433v1 Announce Type: cross Abstract: Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion (non-converged…

  85. arXiv cs.AI TIER_1 English(EN) · Adam Kostka (Warsaw University of Technology), Jaros{\l}aw A. Chudziak (Warsaw University of Technology) ·

    Explainable Belief Harmonization under Dynamic Epistemic Partitions

    arXiv:2607.21210v1 Announce Type: cross Abstract: Existing approaches to multi-agent belief combination have established mature foundations for combining uncertain beliefs under common assumptions: consensus methods use iterative averaging, logic-based methods resolve conflicting…

  86. arXiv cs.CL TIER_1 English(EN) · Ishan S. Kshirsagar ·

    The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning

    arXiv:2607.20952v1 Announce Type: cross Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad the model actively consults during infer…

  87. arXiv cs.AI TIER_1 English(EN) · Wen Ye, Yuxiao Qu, Aviral Kumar, Xuezhe Ma ·

    MIRROR: Learning from the Other View for Multi-Modal Reasoning

    arXiv:2607.21552v1 Announce Type: new Abstract: Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text v…

  88. arXiv cs.AI TIER_1 English(EN) · Bartolomeo Bogliolo ·

    Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog

    arXiv:2607.21412v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approac…

  89. arXiv cs.AI TIER_1 English(EN) · Kilian Rueckschloss (Eberhard Karls Universitaet Tuebingen), Felix Weitkaemper (German University of Digital Science) ·

    How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic Programming

    arXiv:2607.21208v1 Announce Type: new Abstract: Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge supports only conclusions drawn from observations. His theory of causality, however, is developed …

  90. arXiv cs.AI TIER_1 English(EN) · Akihiro Takemura (National Institute of Informatics, Tokyo, Japan), Katsumi Inoue (National Institute of Informatics, Tokyo, Japan) ·

    Differentiable Logic Programming to Mitigate Reasoning Shortcuts in Neurosymbolic Systems

    arXiv:2607.21185v1 Announce Type: new Abstract: Neurosymbolic (NeSy) systems integrate neural networks with logical reasoning to achieve both generalization and interpretability, but recent work has shown they are susceptible to shortcut reasoning behaviors. We propose a novel me…

  91. arXiv cs.AI TIER_1 English(EN) · Sagnik Nath, Edith Aurora Graf, Liang Zhang, Diego Zapata-Rivera ·

    Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

    arXiv:2607.20520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failure…

  92. arXiv cs.AI TIER_1 English(EN) · Sizhe Tang, Guangyu Jiang, Yu Li, Rongqian Chen, Ioannis G. Kevrekidis, Tian Lan ·

    FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts

    arXiv:2607.20500v1 Announce Type: new Abstract: Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in the open world can become ill-posed due to inconsistent conditions, conflicting statements, or …

  93. arXiv cs.CL TIER_1 English(EN) · Zhensheng Jin, Xin Dai, Zhenghao Liu, Chaojun Xiao, Huiyuan Xie, Yu Gu, Ge Yu, Maosong Sun ·

    REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

    arXiv:2607.20833v1 Announce Type: new Abstract: Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge. Existing…

  94. arXiv cs.CL TIER_1 English(EN) · Simone Angarano, Francesco Bertolotti, Federico D'Ambrosio, Michele Resta, Alessandro Rognoni, Nicol\`o Ruggeri, Dario Salvati, Andrea Valenti, Alberto Veneri, Martin Cimmino ·

    Domyn-Small: A European 10B Reasoning Language Model

    arXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an initial pre-training phase on 9 trillion tokens multilingual data, followed by a …

  95. Hugging Face Daily Papers TIER_1 English(EN) ·

    MIRROR: Learning from the Other View for Multi-Modal Reasoning

    Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit diff…

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog

    Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approaches address this gap by coupling neural models w…

  97. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jarosław A. Chudziak ·

    Explainable Belief Harmonization under Dynamic Epistemic Partitions

    Existing approaches to multi-agent belief combination have established mature foundations for combining uncertain beliefs under common assumptions: consensus methods use iterative averaging, logic-based methods resolve conflicting knowledge bases, and epistemic logic analyzes age…

  98. Hugging Face Daily Papers TIER_1 English(EN) ·

    How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic Programming

    Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge supports only conclusions drawn from observations. His theory of causality, however, is developed exclusively within Bayesian networks and causal …

  99. arXiv cs.AI TIER_1 English(EN) · Xinbang Dai, Zheyu Xin, Huikang Hu, Lin Ren, Rihui Jin, Guohui Xiao, Guilin Qi, Kuicai Dong, Zhaocheng Du, Yuyang Zhang ·

    EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization

    arXiv:2607.19962v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switching and reasoning trajectory compression, fail to ma…

  100. arXiv cs.AI TIER_1 English(EN) · Yousef Khan, Luca Gherardini, Marco Maratea, Joel Arrais, Jose Sousa ·

    CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs

    arXiv:2607.19996v1 Announce Type: new Abstract: Machine Learning models are widely used for automating classification tasks by extracting statistical patterns from data. However, their performance deteriorates if the data distribution changes, making them ill-suited to handle unc…

  101. arXiv cs.AI TIER_1 English(EN) · Alexis Fox, Junlin Wang, Paul Rosu, Bhuwan Dhingra ·

    PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning

    arXiv:2607.20064v1 Announce Type: new Abstract: Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks s…

  102. arXiv cs.AI TIER_1 English(EN) · Anmol Kankariya, Sercan \"O. Ar{\i}k ·

    PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

    arXiv:2607.20268v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle…

  103. arXiv cs.AI TIER_1 English(EN) · Hiskias Dingeto ·

    Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

    arXiv:2607.20379v1 Announce Type: new Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claim…

  104. arXiv cs.AI TIER_1 English(EN) · Wael AbdAlmageed ·

    SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data

    arXiv:2607.20402v1 Announce Type: new Abstract: In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledg…

  105. arXiv cs.AI TIER_1 English(EN) · Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean, Maria Liakata ·

    Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

    arXiv:2607.19678v1 Announce Type: cross Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for …

  106. arXiv cs.AI TIER_1 English(EN) · Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li ·

    SLPO: Scaling Latent Reasoning via a Surrogate Policy

    arXiv:2607.19691v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediat…

  107. arXiv cs.AI TIER_1 English(EN) · Miao Li, Alexander Gurung, Irina Saparina, Mirella Lapata ·

    SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles

    arXiv:2509.21028v4 Announce Type: replace Abstract: We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs). Existing long-context datasets with inputs beyond 64K tokens either targe…

  108. arXiv cs.AI TIER_1 English(EN) · Wouter W. L. Nuijten, Bert de Vries ·

    Sophisticated Policies from Epistemic Priors

    arXiv:2607.19518v1 Announce Type: new Abstract: Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its central computational role is simpler: within a planning horizon, it makes active inference …

  109. Hugging Face Daily Papers TIER_1 English(EN) ·

    REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

    Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when evidence is sparse, noisy, or in conflict with parametric knowledge. Existing grounding methods either attach citations after…

  110. arXiv cs.CL TIER_1 English(EN) · Polina Tsvilodub, Fausto Carcassi, Michael Franke ·

    Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    arXiv:2607.18443v1 Announce Type: new Abstract: Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must…

  111. arXiv cs.CL TIER_1 English(EN) · Abir Harrasse, Michael Lan, Hunar Batra, Fateme Hashemi Chaleshtori, Chaithanya Bandi ·

    Reasoning Fine-Tuning Induces Persistent Latent Policy States

    arXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improv…

  112. arXiv cs.AI TIER_1 English(EN) · Liam Swayne ·

    Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

    arXiv:2607.18438v1 Announce Type: cross Abstract: Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The …

  113. arXiv cs.AI TIER_1 English(EN) · Ayhan Suleymanzade, Halil Alperen Gozeten, Michael Bronstein, \.Ismail \.Ilkan Ceylan, Jinwoo Kim ·

    MUX: Continuous Reasoning via Multiplexed Tokens

    arXiv:2607.18264v1 Announce Type: new Abstract: Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoning step conveys only a single subword, and many are …

  114. arXiv cs.AI TIER_1 English(EN) · Dmitrii Kharlapenko, Terry Jingchen Zhang, Arth Singh, Alessandro Stolfo, Arthur Conmy, Mrinmaya Sachan, Zhijing Jin ·

    Fluid Reasoning Representations

    arXiv:2602.04843v2 Announce Type: replace Abstract: Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking. Yet we lack a mechanistic account of how extended thinking changes hidden-state representations ove…

  115. arXiv cs.AI TIER_1 English(EN) · Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang ·

    Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

    arXiv:2607.19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical…

  116. arXiv cs.AI TIER_1 English(EN) · Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi ·

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

    arXiv:2607.19313v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{ze…

  117. arXiv cs.AI TIER_1 English(EN) · Yu Wang, Ming Fan, Xicheng Zhang, Zhiyong Li, Zhihu Wang, Caiyue Xu, Dahai Hu, Ting Liu ·

    DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

    arXiv:2607.19088v1 Announce Type: cross Abstract: Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. W…

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

    Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the recon…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    SLPO: Scaling Latent Reasoning via a Surrogate Policy

    Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent…

  120. Hugging Face Daily Papers TIER_1 English(EN) ·

    DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

    Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Super…

  121. arXiv cs.AI TIER_1 English(EN) · Satyam Kumar, Saurabh Jha ·

    CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

    arXiv:2607.16955v1 Announce Type: cross Abstract: On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to…

  122. arXiv cs.AI TIER_1 English(EN) · Thorir Mar Ingolfsson, Wajeeha Tahir, Anna Tegon, Lionnus Kesting, Gamze \.Islamo\u{g}lu, Luca Benini ·

    Quantizing Recursive Reasoning Models

    arXiv:2607.16237v1 Announce Type: cross Abstract: Recursive reasoning models solve hard puzzles by applying compact, weight-tied blocks over many refinement steps. Because these blocks are reused many times, quantizing them creates a unique dynamical problem: the quantization err…

  123. arXiv cs.AI TIER_1 English(EN) · Brian K Chen ·

    Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

    arXiv:2607.18228v1 Announce Type: new Abstract: To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we chara…

  124. arXiv cs.AI TIER_1 English(EN) · Sheldon Yu, Tong Yu, Xunyi Jiang, Rohan Surana, Gagan Mundada, Sungchul Kim, Lina Yao, Julian McAuley, Junda Wu ·

    Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

    arXiv:2607.18100v1 Announce Type: new Abstract: Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches …

  125. arXiv cs.AI TIER_1 English(EN) · Zehua Cheng, Wei Dai, Jiahao Sun ·

    Constraint-Anchored Reasoning Traces

    arXiv:2607.16727v1 Announce Type: new Abstract: Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning. We find that in state-of-the-art open-sour…

  126. arXiv cs.AI TIER_1 English(EN) · Yanxiao Zhao, Yaqian Li, Zihao Bo, Rinyoichi Takezoe, Haojia Hui, Mo Guang, Lei Ren, Xiaolin Qin, Kaiwen Long ·

    SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs

    arXiv:2509.00930v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit strong general reasoning, yet the community lacks controllable, scalable, and verifiable tools to analyze and improve these abilities. We present SATQuest, a verifier that generates diverse S…

  127. arXiv cs.AI TIER_1 English(EN) · Yosub Shin, Michael Buriek, Boris Sobolev, Pavel Bushuyeu, Vikas Kumar, Haoyang Xu, Samuel Watson, Igor Molybog ·

    Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols

    arXiv:2601.10922v2 Announce Type: replace Abstract: We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constant and the main degree of freedom is the training…

  128. Hugging Face Daily Papers TIER_1 English(EN) ·

    Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that…

  129. arXiv cs.CL TIER_1 English(EN) · Leichao Dong, Dongxu Zhang, Yiding Sun, Qirui Wang, Yuhan Wang, Lin Chen, Jihua Zhu ·

    Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning

    arXiv:2607.15736v1 Announce Type: new Abstract: Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations, repeated self-verification, and detours that do not improve the final answer. Ex…

  130. arXiv cs.AI TIER_1 English(EN) · Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov ·

    Understanding Reasoning from Pretraining to Post-Training

    arXiv:2607.16097v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basi…

  131. arXiv cs.AI TIER_1 English(EN) · Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur ·

    Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

    arXiv:2607.15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 ver…

  132. arXiv cs.AI TIER_1 Italiano(IT) · J. N. Hooker ·

    Logic, Optimization, and Artificial Intelligence

    arXiv:2607.15532v1 Announce Type: new Abstract: Logic and optimization can, in combination, make valuable contributions to rule-based AI. Logic is the obvious medium for encoding a rule base and drawing inferences from it, while optimization provides a powerful technology for com…

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    Uncovering Latent Reasoning Strategies in Language Models

    A language model p_θ(y mid x) trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution. We study the problem of decomposing the response distribution of a given p…

  134. arXiv cs.CL TIER_1 English(EN) · Seungone Kim, Ian Wu, Jinu Lee, Xiang Yue, Seongyun Lee, Mingyeong Moon, Carolin Lawrence, Kiril Gashteovski, Julia Hockenmaier, Graham Neubig, Sean Welleck ·

    Scaling Evaluation-time Compute with Reasoning Models as Evaluators

    arXiv:2503.19877v3 Announce Type: replace Abstract: As language model (LM) outputs get more and more natural, it is becoming more difficult than ever to evaluate their quality. Simultaneously, increasing LMs' "thinking" time through scaling test-time compute has proven an effecti…

  135. arXiv cs.AI TIER_1 English(EN) · Haohua Niu, Xingtong Yu, Yang Liu, Junfeng Fang, Xuanting Xie, Jie Tan, Zhongjian Zhang, Hong Cheng, Yuan Fang ·

    CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning

    arXiv:2607.14114v1 Announce Type: cross Abstract: Graph learning under distribution shift presents a persistent challenge, where models adapt to new graphs with limited or even no supervision. Recent graph--LLM approaches move toward label-efficient prediction by linearizing grap…

  136. arXiv cs.AI TIER_1 English(EN) · Abdullah Shaikh, Zain Naqi, Taha Zahid, Sandesh Kumar, Abdul Samad ·

    HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization

    arXiv:2607.14349v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility. In this paper, we p…

  137. arXiv cs.AI TIER_1 English(EN) · Jungseob Lee, Seungyoon Lee, Suhyune Son, Dongyub Jude Lee, Sungbin Han, Sugyeong Eo, Heuiseok Lim ·

    Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

    arXiv:2607.14552v1 Announce Type: cross Abstract: A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fa…

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    Understanding Reasoning from Pretraining to Post-Training

    Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining ch…

  139. arXiv cs.CL TIER_1 English(EN) · Heuiseok Lim ·

    Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

    A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix shows the generator the gold ans…

  140. arXiv cs.AI TIER_1 English(EN) · Hefeng Zhou, Jinxuan Zhang, Jiong Lou, Yuxin Liu, Chaochao Lu, Jingjing Qu, Jie Li ·

    Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

    arXiv:2607.14049v1 Announce Type: new Abstract: The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically invol…

  141. arXiv cs.CL TIER_1 English(EN) · Abdul Samad ·

    HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization

    While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility. In this paper, we present our system for SemEval-2026 Task 11, which …

  142. arXiv cs.AI TIER_1 English(EN) · Jie Li ·

    Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

    The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically involve re-generating another response that may make …

  143. Hugging Face Daily Papers TIER_1 English(EN) ·

    Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape

    Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback. We study why closed-loop knowledge systems saturate and what external information can m…

  144. arXiv cs.AI TIER_1 English(EN) · Junyu Ren ·

    Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

    arXiv:2607.12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Ver…

  145. arXiv cs.CL TIER_1 English(EN) · Xinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan, Xiaochong Lan, Yifan Li, Yuchen Yan, Han Peng, Zican Dong, Zhenduo Zhang, Tianshu Wang, Xinyu Kong, Zujie Wen, Wayne Xin Zhao, Zhiqiang Zhang, Jun Zhou ·

    Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

    arXiv:2607.12395v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, exist…

  146. arXiv cs.AI TIER_1 English(EN) · Ning Liu ·

    Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

    arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed,…

  147. arXiv cs.AI TIER_1 English(EN) · Dongyue Li, Zhenshuo Zhang, Minxuan Duan, Edgar Dobriban, Hongyang R. Zhang ·

    Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning

    arXiv:2512.01113v2 Announce Type: replace-cross Abstract: Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reasoning abilities, designed for graph neural networks and also for transformer models. …

  148. arXiv cs.AI TIER_1 English(EN) · Julius Steiglechner, Lucas Mahler, Gabriele Lohmann ·

    LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

    arXiv:2607.12733v1 Announce Type: new Abstract: Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce E…

  149. Hugging Face Daily Papers TIER_1 English(EN) ·

    GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

    The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Banglades…

  150. arXiv cs.CL TIER_1 English(EN) · Swastika Kundu ·

    GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

    The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Banglades…

  151. arXiv cs.LG TIER_1 English(EN) · Gabriele Lohmann ·

    LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

    Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cros…

  152. arXiv cs.AI TIER_1 English(EN) · Junyu Ren ·

    Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

    Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Agentic Reasoning), a Lean 4-based tool-call…

  153. arXiv cs.CL TIER_1 English(EN) · Jun Zhou ·

    Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

    Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small mode…

  154. arXiv cs.AI TIER_1 English(EN) · Sabari Iyyappan Duraipandian, Shreya Sanjay Boyane, Manju Nagesh, Jerome Francis, Archana Vaidheeswaran, Kevin Zhu ·

    Interpreting Latent CoT Reasoning as Dynamical Systems

    arXiv:2607.09698v1 Announce Type: new Abstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate traces in the hidden space at each step, unlike explicit- CoT, which follows a sin…

  155. arXiv cs.AI TIER_1 English(EN) · Huan Zhu ·

    Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction

    arXiv:2607.11696v1 Announce Type: new Abstract: Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolat…

  156. arXiv cs.AI TIER_1 English(EN) · Joyjeet Singh ·

    The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

    arXiv:2607.11116v1 Announce Type: cross Abstract: Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary …

  157. arXiv cs.AI TIER_1 English(EN) · Mohan Tang, Sidi Lu ·

    Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers

    arXiv:2602.17993v2 Announce Type: replace-cross Abstract: Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step informs the next. In this work, we adopt the perspective that the reasoning power of Tran…

  158. arXiv cs.AI TIER_1 English(EN) · Kejing Xia, Mingzhe Li, Lixuan Wei, Zhenbang Du, Xiangchi Yuan, Dachuan Shi, Qirui Jin, Wenke Lee ·

    MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models

    arXiv:2603.01331v3 Announce Type: replace-cross Abstract: Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each denoising step solely on the current hard-masked sequence, while intermediate cont…

  159. arXiv cs.AI TIER_1 English(EN) · Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi ·

    Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces

    arXiv:2604.11996v2 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone does not reveal the quality of the reasoning used to produce it. This highlights a …

  160. arXiv cs.AI TIER_1 English(EN) · Mohammed Ehab, Aymane El Gadarri, Vivek F. Farias, Adam Jozefiak, Ciamac C. Moallemi ·

    OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

    arXiv:2607.11089v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning step…

  161. arXiv cs.AI TIER_1 English(EN) · JungMin Yun, JuneHyoung Kwon, YoungBin Kim ·

    CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

    arXiv:2607.10562v1 Announce Type: new Abstract: Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks…

  162. arXiv cs.LG TIER_1 English(EN) · Ning Liu ·

    Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

    Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed, what it suspects, and what it still needs) lives …

  163. Hugging Face Daily Papers TIER_1 English(EN) ·

    Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

    Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small mode…

  164. arXiv cs.AI TIER_1 English(EN) · Huan Zhu ·

    Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction

    Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolation between reasoning stages, so that informatio…

  165. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

    Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned ini…

  166. Hugging Face Daily Papers TIER_1 English(EN) ·

    OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

    Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improvi…

  167. Hugging Face Daily Papers TIER_1 English(EN) ·

    ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

    Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical ag…

  168. arXiv cs.AI TIER_1 English(EN) · Jakob Suchan, Julius Monsen, Salim Baloch, Mehul Bhatt ·

    Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

    arXiv:2607.08136v1 Announce Type: new Abstract: We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the con…

  169. arXiv cs.CL TIER_1 English(EN) · Chris Samarinas, Haw-Shiuan Chang, Hamed Zamani ·

    Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning

    arXiv:2602.23440v4 Announce Type: replace Abstract: Reinforcement learning has emerged as an effective paradigm for training large language models to interleave reasoning with search engine calls. However, existing approaches face a fundamental credit assignment problem: methods …

  170. arXiv cs.AI TIER_1 English(EN) · Jack Hopkins, Dipika Khullar, Fabien Roger ·

    Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

    arXiv:2607.08173v1 Announce Type: new Abstract: Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overt…

  171. arXiv cs.AI TIER_1 English(EN) · Shaoyu Wang, Kaiyue Zhao, Dongliang Wei, Przemys{\l}aw Andrzej Wa{\l}\k{e}ga, Dingmin Wang, Hongming Cai, Pan Hu ·

    Goal-Driven Reasoning in DatalogMTL with Magic Sets

    arXiv:2412.07259v5 Announce Type: replace Abstract: DatalogMTL is a powerful rule-based language for temporal reasoning. Due to its high expressive power and flexible modeling capabilities, it is suitable for a wide range of applications, including tasks from industrial and finan…

  172. arXiv cs.AI TIER_1 English(EN) · Fabien Roger ·

    Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

    Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task ve…

  173. Hugging Face Daily Papers TIER_1 English(EN) ·

    Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

    We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model substrate. Key contributions are: (1) supporting joint optimisation in the continuous latent space through explicit ASP-based …

  174. arXiv cs.CL TIER_1 English(EN) · Xinda Jia, Jinpeng Li, Zezhong Wang, Jingjing Li, Xingshan Zeng, Yasheng Wang, Weinan Zhang, Yong Yu, Weiwen Liu ·

    Fast, Slow, and Tool-augmented Thinking for LLMs: A Review

    arXiv:2508.12265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable progress in reasoning across diverse domains. However, effective reasoning in real-world tasks requires adapting the reasoning strategy to the demands of the problem, ran…

  175. arXiv cs.AI TIER_1 English(EN) · Azwar Abdulsalam, Nishil Patel, Andrew Saxe ·

    RL Post-Training Builds Compositional Reasoning Strategies

    arXiv:2607.07646v1 Announce Type: new Abstract: Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment wher…

  176. arXiv cs.AI TIER_1 English(EN) · Guangzhi Wang, Kai Li, Yinghao Jiao, Zhi Liu ·

    Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning

    arXiv:2511.13726v2 Announce Type: replace-cross Abstract: We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models. The method obtains the final semantic representation by running multiple forward passes of the text embeddi…

  177. arXiv cs.CL TIER_1 English(EN) · Ruilin Tong, Dong Gong ·

    MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning

    arXiv:2607.06974v1 Announce Type: new Abstract: Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience …

  178. arXiv cs.CL TIER_1 English(EN) · Hengyu Jin, Shu Yang, Di Wang ·

    Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories

    arXiv:2607.06648v1 Announce Type: cross Abstract: Latent reasoning methods perform multi-step inference entirely in the model's continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden states raise a question of faithfulness: whether …

  179. arXiv cs.CL TIER_1 English(EN) · Josip Juki\'c, Ivan Titov ·

    Geometric Self-Distillation for Reasoning Generalization

    arXiv:2607.06855v1 Announce Type: cross Abstract: On-policy distillation is a practical post-training recipe for large language models, supplying dense teacher supervision on the student's own trajectories. In privileged-context self-distillation, teacher and student are the same…

  180. arXiv cs.AI TIER_1 English(EN) · Dmitry Beresnev, Vladimir Makharev, Roman Khalikov, Ivan Oseledets, Petr Anokhin ·

    Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

    arXiv:2607.07492v1 Announce Type: new Abstract: Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent…

  181. arXiv cs.AI TIER_1 English(EN) · Andrew Saxe ·

    RL Post-Training Builds Compositional Reasoning Strategies

    Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and ever…

  182. arXiv cs.AI TIER_1 English(EN) · Petr Anokhin ·

    Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

    Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent, a training and inference framework inspired by…

  183. arXiv cs.AI TIER_1 English(EN) · He Liu, Changtao Miao, Xinjie Yang, Tianle Song, Yin Wu, Junchi Chen, Bintao He, Xinyuan Zhang, Bo Zhang, Shi Yan, Wei Lu, Wei Wang, Danyang Xu, Jiansheng Cai, Zhe Li ·

    DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

    arXiv:2607.06326v1 Announce Type: new Abstract: Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. Existing guardrails face a practical trade-off betwee…

  184. arXiv cs.AI TIER_1 English(EN) · Yang Liu, Zhaokai Luo, Huayi Jin, Ruozhou He, Chenchen Hong, Zhiyong Wang, Yifei Liu, Yunfei Gu, Chentao Wu, Junhao Hu ·

    Akashic: A Low-Overhead LLM Inference Service with MemAttention

    arXiv:2607.05708v1 Announce Type: new Abstract: Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts in…

  185. arXiv cs.CL TIER_1 English(EN) · Dong Gong ·

    MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning

    Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Exi…

  186. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

    Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for go…

  187. arXiv cs.CL TIER_1 English(EN) · Ivan Titov ·

    Geometric Self-Distillation for Reasoning Generalization

    On-policy distillation is a practical post-training recipe for large language models, supplying dense teacher supervision on the student's own trajectories. In privileged-context self-distillation, teacher and student are the same model conditioned on the same prefix, but the tea…

  188. arXiv cs.CL TIER_1 English(EN) · Di Wang ·

    Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories

    Latent reasoning methods perform multi-step inference entirely in the model's continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden states raise a question of faithfulness: whether these latent reasoning steps causally drive the fi…

  189. arXiv cs.AI TIER_1 English(EN) · Zhe Li ·

    DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

    Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. Existing guardrails face a practical trade-off between lightweight classification-based models, which…

  190. arXiv cs.AI TIER_1 English(EN) · Hehai Lin, Shilei Cao, Sudong Wang, Haotian Wu, Minzhi Li, Linyi Yang, Juepeng Zheng, Chengwei Qin ·

    Interactive Learning for LLM Reasoning

    arXiv:2509.26306v5 Announce Type: replace Abstract: Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multiple Large Language Models (LLMs), thereby constructing stronger multi-agent systems (MAS). …

  191. arXiv cs.AI TIER_1 English(EN) · Jingchu Wang, Bingbing Xu, Yige Yuan, Dan Zhang, Bin Xie, Xiaoqian Sun, Huawei Shen ·

    R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning

    arXiv:2601.11960v3 Announce Type: replace-cross Abstract: Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincide with the one producing inference responses. We argue that this is a misleading in…

  192. arXiv cs.AI TIER_1 English(EN) · Mengqi Li, Lei Zhao, Anthony Man-Cho So, Ruoyu Sun, Xiao Li ·

    A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

    arXiv:2510.18814v4 Announce Type: replace-cross Abstract: Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolving Post-Training (SePT), a simple post-tra…

  193. arXiv cs.CL TIER_1 English(EN) · Wei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao, Blake JianHang Chen, Ziqian Lin, Alec Go, Yu Meng ·

    Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

    arXiv:2602.13517v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). However, recent findings suggest that raw token counts are unreliable proxies for rea…

  194. Hugging Face Daily Papers TIER_1 English(EN) ·

    Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming

    Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning …

  195. Hugging Face Daily Papers TIER_1 English(EN) ·

    RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules

    RuleChef utilizes large language models to generate and iteratively improve executable rules for NLP tasks through example-based learning and human feedback, resulting in fast and inspectable rule systems.

  196. arXiv cs.CV TIER_1 English(EN) · Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao ·

    ChronoVision: Temporal Reasoning via Latent State Reconstruction

    arXiv:2608.05631v1 Announce Type: new Abstract: Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reas…

  197. arXiv stat.ML TIER_1 English(EN) · Cai Zhou, Zekai Wang, Menghua Wu, Qianyu Julie Zhu, Flora C. Shi, Chenyu Wang, Ashia Wilson, Tommi Jaakkola, Stephen Bates ·

    Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning

    arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-tra…

  198. arXiv stat.ML TIER_1 English(EN) · Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho ·

    When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

    arXiv:2608.03506v1 Announce Type: cross Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple vali…

  199. arXiv cs.CV TIER_1 English(EN) · Jiaxuan Kang, Siyu Chen, Mingda Li, Mingjie Liu, Tianyue Wang, Zhaoyang Wei, Yongheng Zhang, Yanchao Hao, Zheng Wei ·

    LUT: Latent Utility Training for Visual Reasoning

    arXiv:2608.00743v1 Announce Type: new Abstract: Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden-space computation before answering, but they often …

  200. arXiv cs.CV TIER_1 English(EN) · Shoutai Zhu, Tianyang Xu, Bin Sun, Mingyuan Xu, Yu Liu, Qinzhen Guo ·

    OPLD: On-Policy Latent Distillation for Multimodal Reasoning

    arXiv:2607.28154v1 Announce Type: new Abstract: Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces a…

  201. arXiv stat.ML TIER_1 English(EN) · Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Weijie J. Su, Edgar Dobriban ·

    Statistical Early Stopping for Reasoning Models

    arXiv:2602.13935v2 Announce Type: replace-cross Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introd…

  202. arXiv stat.ML TIER_1 English(EN) · Zihan Dong, Zhixian Zhang, Yang Zhou, Can Jin, Ruijia Wu, Linjun Zhang ·

    Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

    arXiv:2602.03061v2 Announce Type: replace-cross Abstract: Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms. On difficult problems,…

  203. LessWrong (AI tag) TIER_1 English(EN) · Chandram Dutta ·

    The Termination Circuit (how reasoning models stop thinking).

    <p><span>Reasoning models since the dawn of o1 and R1 have a tendency to overthink. Despite a lot of work on early-exit methods and steering, open-weight and smaller reasoning models still produce long chains of thought before they answer. I worked on discovering how much of that…

  204. AWS Machine Learning Blog TIER_1 English(EN) · Rushil Anirudh ·

    Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

    In this post, we explore an idea for generating thinking tokens for datasets that lack reasoning traces in SFT customization. We first examine the reasoning suppression problem, then introduce Self-Distilled Reasoning (SDR), validate it across three benchmarks, and provide practi…

  205. Sequoia Capital TIER_1 English(EN) · amoore ·

    Partnering with Etched: Building the Inference Machine

    <p>The post <a href="https://sequoiacap.com/article/partnering-with-etched-building-the-inference-machine/">Partnering with Etched: Building the Inference Machine</a> appeared first on <a href="https://sequoiacap.com">Sequoia Capital</a>.</p>

  206. Medium — fine-tuning tag TIER_1 English(EN) · Stancho Stanchev ·

    Fine-Tuning Qwen3–4B vs. SmolLM3–3B on the Same Math Reasoning Recipe

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@stanchoz3/fine-tuning-qwen3-4b-vs-smollm3-3b-on-the-same-math-reasoning-recipe-9fad07957de3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2600/0*m0hETYuphBsrKtcb…

  207. Medium — Claude tag TIER_1 English(EN) · Nadeem Khan(NK) ·

    Chain-of-Thought and Scratchpad Reasoning: The Model Thinking Out Loud

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://nadeem4-nk13.medium.com/chain-of-thought-and-scratchpad-reasoning-the-model-thinking-out-loud-d998a5a1860d?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/0*4E12i9OTIivI4pFp" …

  208. Medium — fine-tuning tag TIER_1 Bahasa(ID) · Angga Yulian Adi Pradana ·

    GRPO Fine-Tuning LLM for Training Reasoning Models

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anggapradanaa/grpo-fine-tuning-llm-untuk-melatih-reasoning-model-6adbf065c48a?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1734/1*xyCJt-bhgqHvLf3rS_197g.png" wi…

  209. dev.to — LLM tag TIER_1 English(EN) · Hunter G ·

    The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

    <h1> The Inference Engineering Masterclass — Philip Kiely &amp; Ali Taha, Baseten </h1> <p>The same open-weights model, served by different providers, can differ by 4x to 10x in speed. So "which model" only answers half the question. The other half is whose inference.</p> <p>Base…

  210. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    ReflectRL: Turning Failed LLM Reasoning into Training Signals

    <p><em>ReflectRL introduces a novel framework that utilizes 'Golden Negative Trajectories'—failed reasoning attempts by expert models—to improve LLM performance. By treating these failures as opportunities for reflection rather than discarding them, the method enhances reasoning …

  211. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Reasoning errors in LLMs aren't just random noise; they are anchored to specific directions within the residual stream. Identifying these high-dimensional vecto

    Reasoning errors in LLMs aren't just random noise; they are anchored to specific directions within the residual stream. Identifying these high-dimensional vectors allows for targeted interventions, moving us from black-box prompting to surgical model steering. # LLMs # AI

  212. r/LocalLLaMA TIER_1 English(EN) · /u/jwdeaver ·

    LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vhz2rc/labyrinthbench_a_localfocused_judgefree_llm/"> <img alt="LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks." src="ht…

  213. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    From Prompting to Reasoning: How ToolArtist Unifies Multi-Step Logic and Image Generation

    <h1> From Prompting to Reasoning: How ToolArtist Unifies Multi-Step Logic and Image Generation </h1> <p>Text-to-image (T2I) systems have reached a level of visual fidelity that was difficult to imagine only a few years ago. However, even the most advanced diffusion models and aut…

  214. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Claude Opus 5: What the ARC-AGI-3 Leap Actually Tells Us About Reasoning Progress

    <h1> Claude Opus 5: What the ARC-AGI-3 Leap Actually Tells Us About Reasoning Progress </h1> <p>Anthropic released <a href="https://www.anthropic.com/news/claude-opus-5" rel="noopener noreferrer">Claude Opus 5</a> on July 24, 2026, and the headline number is hard to ignore: a 30.…

  215. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    Chain-of-Draft: keep the reasoning, drop the narration, and cut ~80% of your reasoning tokens

    <p>Chain-of-Thought reliably lifts reasoning accuracy by making a model write its intermediate steps down instead of leaping to an answer. But look at a CoT trace closely and you'll notice something: on a simple arithmetic problem the model emits a paragraph where the actual work…

  216. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    Contrastive Chain-of-Thought: pair each correct rationale with a labelled wrong one, teaching the model the reasoning boundary

    <p>Plain chain-of-thought hands the model a few <em>correct</em> worked examples and hopes it generalizes the right steps. It's genuinely strong — but I kept watching it walk into the same trap on multi-step arithmetic: adding when it should subtract, jumbling the step order, gra…

  217. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    SIGMA transforms deterministic AI reasoning into transparent, traceable language without compromising reproducibility or governance. Discover NEUROVATIC's sover

    SIGMA transforms deterministic AI reasoning into transparent, traceable language without compromising reproducibility or governance. Discover NEUROVATIC's sovereign language intelligence architecture: https:// neurovatic.ai/sigma # AI # SovereignAI # SIGMA

  218. dev.to — LLM tag TIER_1 English(EN) · Daniel Dong ·

    Kimi K3: The reasoning model that doesn't have a "think" toggle

    <p>A reasoning model that doesn't make you turn on "think mode."</p> <p>Kimi K3 thinks before every answer — no toggle, no config, no prompt engineering.</p> <p>1M context window. Drop a codebase. Get a review. No chunks required.<br /> </p> <div class="highlight js-code-highligh…

  219. r/LocalLLaMA TIER_1 English(EN) · /u/hellajacked ·

    MindControl - llama.cpp fork to guide the reasoning process via injection during sampling

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v3ms3c/mindcontrol_llamacpp_fork_to_guide_the_reasoning/"> <img alt="MindControl - llama.cpp fork to guide the reasoning process via injection during sampling" src="https://preview.redd.it/yo4avmy3dteh1.png?w…

  220. dev.to — LLM tag TIER_1 English(EN) · Satvik Mishra ·

    KNOT — an experimental reasoning architecture that argues with itself (first prototype showcase)

    <p>Meet the <em>knot</em>. I spent a month making him argue with himself.<br /> <a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.co…

  221. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Verified Agentic Reasoning is the necessary pivot from probabilistic generation to verifiable output. By anchoring LLM inferences in tool-attested kernel proofs

    Verified Agentic Reasoning is the necessary pivot from probabilistic generation to verifiable output. By anchoring LLM inferences in tool-attested kernel proofs, we move beyond "plausible" text to functional accuracy, essential for high-stakes empirical science. # AI # LLMs (1/2)

  222. dev.to — LLM tag TIER_1 English(EN) · Tech Signal Daily ·

    How Logically Inconsistent Prompts Can Turn Reasoning Models Into a DoS Problem

    <p>Reasoning models are supposed to be better at hard tasks because they “think” before answering. In practice, that extra step-by-step process creates a new attack surface: if you can push the model into overthinking, you can make it spend far more tokens than normal and slow it…

  223. r/LocalLLaMA TIER_1 English(EN) · /u/marcodsn ·

    [Study/Models] Flint: Compressing Reasoning Without Breaking It

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uv9o2u/studymodels_flint_compressing_reasoning_without/"> <img alt="[Study/Models] Flint: Compressing Reasoning Without Breaking It" src="https://external-preview.redd.it/oW5M6NGWwIm7UmrIEq6LxCtz9PZYMnGJKCXJx…

  224. dev.to — LLM tag TIER_1 English(EN) · jackma ·

    Building a Study App Around LLM Reasoning

    <p><strong>Building a Study App Around LLM Reasoning</strong></p> <p>When building an AI study app, it is easy to focus on the output: the answer, the final number, the completed explanation.</p> <p>But the more interesting product question is what happens before and around that …

  225. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Thinking Machines releases Inkling Small: Open-weights reasoning model with less than a third of the parameters of its predecessor, but better benchmark results

    Thinking Machines veröffentlicht Inkling Small: Open-Weights-Reasoning-Modell mit unter einem Drittel der Parameter des Vorgängers, aber besseren Benchmark-Ergebnissen. Effizienzgewinn durch Architektur statt reiner Skalierung. https:// the-decoder.de/thinking-machin es-setzt-mit…

  226. Mastodon — mastodon.social TIER_1 English(EN) · strike007 ·

    The shift toward tool-attested proofs exposes the fragility of current black-box models. Until we bridge the gap between stochastic reasoning and formal verific

    The shift toward tool-attested proofs exposes the fragility of current black-box models. Until we bridge the gap between stochastic reasoning and formal verification, every agentic workflow remains a liability in environments requiring audit-ready certainty. # AI # Safety (2/2)

  227. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    Neurosymbolic AI paper fuses logic rules with energy models A new arXiv preprint fuses Answer Set Programming with energy-based models for end-to-end AI reasoni

    Neurosymbolic AI paper fuses logic rules with energy models A new arXiv preprint fuses Answer Set Programming with energy-based models for end-to-end AI reasoning, tested on visual QA and multi-object tracking benchmarks https://www. notatechguy.com/neurosymbolic- ai-paper-fuses-…

  228. r/singularity TIER_2 English(EN) · /u/yogthos ·

    Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1uyp9xe/ringzero_scaling_zero_rl_to_a_trillion_parameters/"> <img alt="Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning" src="https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnC…