PulseAugur
EN
LIVE 09:55:20

New research tackles LLM faithfulness, reasoning, and evaluation

Recent research explores methods to improve the reliability and evaluation of large language models (LLMs). Several papers introduce frameworks and benchmarks for assessing LLM faithfulness to evidence, detecting inconsistencies in model specifications, and evaluating reasoning capabilities. One study proposes VERITYGATE to check LLM narrations against structured evidence, revealing significant failure rates in current models like GPT-4o-mini and Claude Sonnet 4.6. Another paper introduces VeriSpec, which uses an LLM as a verifier to find inconsistencies in model specifications, successfully identifying issues in the OpenAI Model Spec. Additional research focuses on creating ontology-grounded benchmarks for scientific AI, evaluating generalization through stability rather than just accuracy, and understanding how LLMs transfer mathematical knowledge based on reasoning approach rather than topic. AI

IMPACT These studies aim to improve LLM reliability, reasoning accuracy, and evaluation methodologies, potentially leading to more trustworthy and capable AI systems.

RANK_REASON Multiple research papers published on arXiv detailing new frameworks, benchmarks, and evaluation methods for LLMs.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 96 sources. How we write summaries →

New research tackles LLM faithfulness, reasoning, and evaluation

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Multiple research papers published on arXiv detailing new frameworks, benchmarks, and evaluation methods for LLMs.
Source corroboration
96 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
10 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+47 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [96]

  1. arXiv cs.CL TIER_1 English(EN) · Guang Yang, Homa Hosseinmardi, Fengchen Liu, Amir Ghasemian ·

    Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding

    arXiv:2610.08840v1 Announce Type: new Abstract: Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it …

  2. arXiv cs.LG TIER_1 English(EN) · Harry Mayne, Justin Singh Kang, Dewi Gould, Kannan Ramchandran, Adam Mahdi, Noah Y. Siegel ·

    A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

    arXiv:2602.02639v2 Announce Type: replace-cross Abstract: LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, ty…

  3. arXiv cs.LG TIER_1 English(EN) · Wenqi Li, Bin Liu, Mindi Ruan, Chuanbo Hu, Minglei Yin, Xin Li ·

    Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions

    arXiv:2610.09229v1 Announce Type: new Abstract: LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We …

  4. arXiv cs.CL TIER_1 English(EN) · Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu ·

    SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

    arXiv:2609.33181v2 Announce Type: replace-cross Abstract: Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or…

  5. arXiv cs.CL TIER_1 English(EN) · Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah ·

    Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation

    arXiv:2609.35308v2 Announce Type: replace Abstract: Large language models treat conversation history as unverified context, so false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination prot…

  6. arXiv cs.CL TIER_1 English(EN) · Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu, Ning Song, Yongqi Li, Wenjie Li ·

    EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

    arXiv:2610.10533v1 Announce Type: new Abstract: Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this archite…

  7. arXiv cs.LG TIER_1 English(EN) · Chen Chen, Dongjie Wang, Mei Liu, Zijun Yao ·

    Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence

    arXiv:2610.07739v1 Announce Type: new Abstract: Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interaction…

  8. arXiv cs.AI TIER_1 English(EN) · Maoqi Liu, Quan Fang, Yufei He ·

    CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLMs via Knowledge Consolidation and Decoupling

    arXiv:2610.08312v1 Announce Type: new Abstract: Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-L…

  9. arXiv cs.AI TIER_1 English(EN) · Bach Nguyen, Zhaonan Li, Mau Son Nguyen, Sanika Chavan, Nilay Kumar, Hong Anh Nguyen, Khoa Vo, Ben Zhou ·

    How Well Do LLMs Reason with Noisy Evidence? An Active Visual Reasoning Benchmark

    arXiv:2610.07751v1 Announce Type: new Abstract: Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environm…

  10. arXiv cs.CL TIER_1 English(EN) · Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas-Hinostroza, Mattia Nee, Eliot Krzystof Jones, Ir\`ene Girard, David Mach, Anastasia Stasenko, Ivan P. Yamshchikov ·

    Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

    arXiv:2506.01732v4 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which rai…

  11. arXiv cs.AI TIER_1 Français(FR) · Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban, Stuart Burrell ·

    Latent space bias directions in LLMs capture confidence, not fairness

    arXiv:2610.08559v1 Announce Type: cross Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model perform…

  12. arXiv cs.CL TIER_1 English(EN) · Yiqi Liu, Joseph James, Yang Wang, Kun Zhao, Chenghao Xiao, Chenghua Lin ·

    JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

    arXiv:2610.07109v1 Announce Type: new Abstract: When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights…

  13. arXiv cs.CL TIER_1 English(EN) · Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Borui Jiang, Bo Han ·

    Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective

    arXiv:2610.07894v1 Announce Type: new Abstract: Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically ass…

  14. arXiv cs.CL TIER_1 English(EN) · Kyojun Choo, Minsoo Song, Yunju Kang, Chanjun Park ·

    Structured but Silent: Probing Capability Requirements in LLM Hidden States

    arXiv:2610.08018v1 Announce Type: new Abstract: Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper,…

  15. arXiv cs.AI TIER_1 English(EN) · Hashmath Shaik, Gnaneswar Villuri, Alex Doboli ·

    Will the Judge Flip? Predicting Position-Sensitive LLM Judgments from Residual Stream Activations

    arXiv:2610.07115v1 Announce Type: cross Abstract: The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whe…

  16. arXiv cs.AI TIER_1 English(EN) · Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang ·

    Sherpa: Teaching LLMs to Teach Adaptively

    arXiv:2610.08778v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, pre…

  17. arXiv cs.AI TIER_1 English(EN) · Bingnan Xiao, Chenhao Yang, Bingcong Li, Wei Ni, Xin Wang ·

    Adaptive Power Sampling for LLM Reasoning

    arXiv:2610.08563v1 Announce Type: new Abstract: Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the …

  18. arXiv cs.CL TIER_1 English(EN) · Sofia Torres, Gabriel Almeida, Carter Adams, Camila Rocha ·

    When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO

    arXiv:2610.06861v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \e…

  19. arXiv cs.LG TIER_1 English(EN) · Yuhwan Jeong, Jinnyeong Yang, Kuk-Jin Yoon ·

    Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions

    arXiv:2610.08129v1 Announce Type: new Abstract: Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiv…

  20. arXiv cs.AI TIER_1 English(EN) · Hans Schabert, Christoph Peters ·

    One Step at a Time: Trading LLM Autonomy for Process Predictability

    arXiv:2610.07817v1 Announce Type: cross Abstract: Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictabil…

  21. arXiv cs.AI TIER_1 English(EN) · Dani Roytburg, Daphne Ippolito ·

    Disentangling Models from Personas in Heterogeneous LLM Simulations

    arXiv:2610.07535v1 Announce Type: cross Abstract: Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show …

  22. arXiv cs.AI TIER_1 English(EN) · Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham, Mohit Bansal, Suhyun Kim ·

    Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

    arXiv:2610.07532v1 Announce Type: cross Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, e…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

    Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple…

  24. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Paul Mutawe ·

    From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM-Generated Customer Intents

    Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effective…

  25. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Daphne Ippolito ·

    Disentangling Models from Personas in Heterogeneous LLM Simulations

    Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network p…

  26. Hugging Face Daily Papers TIER_1 English(EN) ·

    Sherpa: Teaching LLMs to Teach Adaptively

    Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria…

  27. arXiv cs.AI TIER_1 English(EN) · Hongji Pu, Yilun Zhao, Wenpeng Yin ·

    PAPER2LLM++: Continual Self-Evolution of LLMs from Research Papers

    arXiv:2610.02793v1 Announce Type: new Abstract: Research on LLMs continually uncovers model limitations, their causes, and potential solutions. Yet these human discoveries remain largely disconnected from model evolution: an LLM does not automatically learn from new research abou…

  28. arXiv cs.AI TIER_1 English(EN) · Ritika Rekhi, Bing Zhang, Md Arman Islam, Jaya Krishna Pasham, Yeswanth Chitturi, Akshay Paramesha, Isha Valiveti, Asif Imran, Bekir Turkkan, Tevfik Kosar ·

    Improving the Energy-Efficiency of the Code Generated by LLMs through Effective Prompting

    arXiv:2610.02571v1 Announce Type: cross Abstract: As AI-assisted programming becomes increasingly mainstream, the environmental impact of AI-generated software has emerged as an important consideration. This motivates evaluating LLM-generated code beyond functional correctness by…

  29. arXiv cs.AI TIER_1 English(EN) · Katherine Van Koevering, Jon Kleinberg ·

    Heads, Tails, and AI Fails: LLMs, Randomness, and Human Judgments

    arXiv:2406.00092v2 Announce Type: replace Abstract: Randomness is central to human cognition and to many applications in which large language models are deployed, yet probabilistic token generation does not imply that LLMs can produce unbiased random sequences. We study how conte…

  30. arXiv cs.AI TIER_1 English(EN) · Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan, Boyu Fang, Yuchen Liu, Jiayi Liu, Hanchao Yu, Qi Guo, Jianyu Wang, Fei Liu, Xiangjun Fan ·

    GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

    arXiv:2603.29112v2 Announce Type: replace Abstract: We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item pre…

  31. arXiv cs.AI TIER_1 English(EN) · Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang ·

    Rhetorical Questions in LLM Representations: A Linear Probing Study

    arXiv:2604.14128v3 Announce Type: replace-cross Abstract: Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unclear. We analyze rhetorical questions in LLM representations using lin…

  32. arXiv cs.CL TIER_1 English(EN) · Siyu Wang, Yifan Wang, Yuecheng He ·

    MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

    arXiv:2610.03080v1 Announce Type: cross Abstract: Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader descr…

  33. arXiv cs.CL TIER_1 English(EN) · Taiqiang Wu, Yuxin Cheng, Chenchen Ding, Runming Yang, Xincheng Feng, Wenyong Zhou, Zhengwu Liu, Ngai Wong ·

    Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality

    arXiv:2603.13725v2 Announce Type: replace Abstract: Memristor-based analog compute-in-memory (CIM) architectures provide a promising substrate for the efficient deployment of Large Language Models (LLMs), owing to superior energy efficiency and computational density. However, the…

  34. arXiv cs.LG TIER_1 English(EN) · Seyedarmin Azizi, Erfan Baghaei Potraghloo, Minoo Ahmadi, Souvik Kundu, Massoud Pedram ·

    Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning

    arXiv:2602.10273v3 Announce Type: replace-cross Abstract: Reasoning ability in large language models is often attributed to \emph{distribution sharpening}: concentrating output probability on high-likelihood sequences. Recent works show that this sharpening effect can be obtained…

  35. arXiv cs.LG TIER_1 English(EN) · Dario Vajda ·

    Teaching LLMs to See Graphs: Unifying Text and Structural Reasoning

    arXiv:2605.10247v2 Announce Type: replace Abstract: Applying Large Language Models (LLMs) to graph-structured data usually involves multi-step pipelines in which textual node attributes are compressed into single tokens and further processed by GNNs, discarding most of their sema…

  36. arXiv cs.AI TIER_1 English(EN) · Seoyeon Ye, Gayoung Kim, Jiyoung Hong, Sookyung Kim, Hyunsoo Cho ·

    LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

    arXiv:2610.02902v1 Announce Type: new Abstract: Current analyses of LLMs' parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct …

  37. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

    Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this…

  38. arXiv cs.LG TIER_1 English(EN) · Sachin Gupta ·

    VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence

    arXiv:2610.00833v1 Announce Type: cross Abstract: Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify e…

  39. arXiv cs.AI TIER_1 English(EN) · Ziyang Yu, Liang Zhao ·

    Distilling LLM Reasoning into Graph of Concept Predictors

    arXiv:2602.03006v3 Announce Type: replace Abstract: Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active distillation reduces these costs by querying an LLM oracle to train small discrim…

  40. arXiv cs.AI TIER_1 English(EN) · Vishvesh Bhat ·

    Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs

    arXiv:2608.28421v2 Announce Type: replace Abstract: Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by ste…

  41. arXiv cs.AI TIER_1 English(EN) · Hongbo Chen, Guohua Lu, Ting Dang, Hong Jia ·

    Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning

    arXiv:2609.33149v2 Announce Type: replace Abstract: A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze sol…

  42. arXiv cs.LG TIER_1 English(EN) · Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang ·

    RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

    arXiv:2609.39801v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ n…

  43. arXiv cs.AI TIER_1 English(EN) · Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li ·

    Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

    arXiv:2610.00296v1 Announce Type: cross Abstract: Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are use…

  44. arXiv cs.AI TIER_1 English(EN) · Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler ·

    Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

    arXiv:2610.00682v1 Announce Type: new Abstract: Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, pr…

  45. arXiv cs.AI TIER_1 English(EN) · Khawaja Murad ul Hassan, Mehran Ebrahimi ·

    Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

    arXiv:2610.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-c…

  46. arXiv cs.AI TIER_1 English(EN) · Sajad Goudarzi, Samaneh Zamanifard, Seyed Amin Seyed Haeri, Moloud Nasiri, Hamed Rahimian ·

    Mathematical Transfer in LLMs Follows Reasoning Approach More Than Topic

    arXiv:2610.00331v1 Announce Type: new Abstract: When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning approach: worked solutions that share a solution method with the …

  47. arXiv cs.AI TIER_1 English(EN) · Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar ·

    The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

    arXiv:2610.00332v1 Announce Type: cross Abstract: Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewa…

  48. arXiv cs.AI TIER_1 English(EN) · Taye Akinrele, Noorbakhsh Amiri Golilarz, Subash Neupane, Sudip Mittal, Shahram Rahimi ·

    Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning

    arXiv:2610.00562v1 Announce Type: cross Abstract: Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts o…

  49. arXiv cs.AI TIER_1 English(EN) · Tae-Eun Song ·

    When Does a Second Model Help? Cross-Model Review in LLM Verification

    arXiv:2610.01471v1 Announce Type: cross Abstract: Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varie…

  50. arXiv cs.AI TIER_1 English(EN) · Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi ·

    Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

    arXiv:2610.01428v1 Announce Type: cross Abstract: Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggre…

  51. arXiv cs.AI TIER_1 English(EN) · Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang ·

    Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning

    arXiv:2610.01847v1 Announce Type: cross Abstract: Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable …

  52. arXiv cs.AI TIER_1 English(EN) · Olivia Macmillan-Scott, Mirco Musolesi ·

    ContextAdapt: Evaluating Contextual Adaptation and Value Alignment in LLMs

    arXiv:2609.38260v1 Announce Type: cross Abstract: Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision …

  53. arXiv cs.AI TIER_1 English(EN) · Guangsheng Yu, Litianyi Zhang, Qin Wang, Xu Wang, Mingyuan Li, Shaoxiong Ji, Ren Ping Liu, Massimo Piccardi ·

    Diversity Combining for Multi-Path LLM Reasoning

    arXiv:2609.38829v1 Announce Type: new Abstract: Multi-path reasoning methods such as self-consistency (SC) sample $K$ reasoning paths and choose the most frequent answer. However, their gains quickly plateau as $K$ increases, and existing methods do not predict when this saturati…

  54. arXiv cs.AI TIER_1 English(EN) · Qijia He, Yu Huang, Yuan Cheng, Yuxin Chen, Yingbin Liang ·

    Provable Test-Time Scaling for Beam Search in LLM Reasoning

    arXiv:2609.38672v1 Announce Type: cross Abstract: Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning effici…

  55. arXiv cs.AI TIER_1 English(EN) · Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi ·

    Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

    arXiv:2609.38972v1 Announce Type: cross Abstract: Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be chan…

  56. arXiv cs.AI TIER_1 English(EN) · Yuliana Shakhvalieva, Dmitrii Kharchev, Viacheslav Bezrukov, Inessa Fedorova, Dmitry Bocharov, Ivan Oseledets, Valerii Ternovskii ·

    What Limits Recursive Reasoning Models: Optimization, Architecture and Test-Time Scaling

    arXiv:2609.39967v1 Announce Type: cross Abstract: Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are …

  57. arXiv cs.CL TIER_1 English(EN) · Bohan Zhang (Southeast University), Linan Yue (Southeast University), Weibo Gao (Hong Kong Polytechnic University), Pengyu Chen (Southeast University), Hong Guo (Southeast University), Yanqi Hao (ZTE Corporation) ·

    Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models

    arXiv:2609.39346v1 Announce Type: new Abstract: Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability…

  58. arXiv cs.CL TIER_1 English(EN) · Michaela Regneri, Nina Scheller, S\"oren Laue ·

    Compact Language, Complex Model Shifts: How and Where Ambiguity and Underspecification Affect LLMs

    arXiv:2609.39572v1 Announce Type: new Abstract: We analyze how lexical ambiguity and underspecification affect language model training. We create artificial homonyms and artificial hypernyms as pseudowords and analyze the generative performance of language models as they are trai…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

    Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or …

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    What Limits Recursive Reasoning Models: Optimization, Architecture and Test-Time Scaling

    Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call …

  61. arXiv cs.LG TIER_1 English(EN) · Shenghong Dai, Shiva Kumar Pentyala, Yingchi Liu, Shubham Mehrotra, Suman Banerjee, James Zhu, Bin Bi, Sitaram Asur, Phil Mui ·

    Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs

    arXiv:2609.36115v1 Announce Type: new Abstract: Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, mu…

  62. arXiv cs.CL TIER_1 English(EN) · Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman, Natalie Harris, Shashir Reddy, Romina Stella, Ariel Goldstein, Marian Croak, Yossi Matias, Amir Feder ·

    Evaluating Alignment of Behavioral Dispositions in LLMs

    arXiv:2602.11328v2 Announce Type: replace Abstract: As people turn to LLMs for social advice, understanding their behavior in such contexts becomes essential. In this work, we focus on behavioral dispositions: the underlying tendencies that shape responses in social contexts. We …

  63. arXiv cs.CL TIER_1 English(EN) · Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni ·

    Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

    arXiv:2609.37915v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher …

  64. arXiv cs.CL TIER_1 English(EN) · Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye ·

    LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

    arXiv:2609.38137v1 Announce Type: new Abstract: Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated acc…

  65. arXiv cs.CL TIER_1 English(EN) · Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko ·

    It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

    arXiv:2609.37891v1 Announce Type: new Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have…

  66. arXiv cs.CL TIER_1 English(EN) · Thomas Reiter, Christoph Kern, Fedor Miasnikov, Sofiia Nikolenko, Rob Chew, Stephanie Eckman, Frauke Kreuter ·

    Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation

    arXiv:2609.35824v1 Announce Type: new Abstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets lab…

  67. arXiv cs.AI TIER_1 English(EN) · Ali Mohammadi Esfahani, Nafiseh Kahani, Samuel A. Ajila ·

    Complexity-Aware Evaluation of LLM Comprehension

    arXiv:2609.37405v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate…

  68. arXiv cs.AI TIER_1 English(EN) · Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls ·

    Causal and Interpretable Structures in LLM Compositional Tasks

    arXiv:2609.35970v1 Announce Type: cross Abstract: Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We s…

  69. arXiv cs.AI TIER_1 English(EN) · Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang ·

    Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

    arXiv:2609.38070v1 Announce Type: new Abstract: As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current meth…

  70. arXiv cs.AI TIER_1 English(EN) · Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui, Shuyi Miao, Pengyang Shao, Yu Zheng, Fei Shen, Tat-Seng Chua ·

    actr: aligning thoughts and responses for multilingual safety in reasoning llms

    arXiv:2609.37054v1 Announce Type: new Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe r…

  71. arXiv cs.AI TIER_1 English(EN) · Fahd Seddik, Fatemeh Fard ·

    Principled Thoughts for Latent Recursive LLM Systems

    arXiv:2609.36159v1 Announce Type: new Abstract: Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final deco…

  72. arXiv cs.LG TIER_1 English(EN) · Qihao Wen, Jiahao Wang, Yang Nan, Pengfei He, Ravi Tandon, Han Xu ·

    Embedding Perturbation may Better Reflect Intermediate-Step Uncertainty in LLM Reasoning

    arXiv:2602.02427v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading outputs. For responsible LLM applications, uncertainty quantification techniques ar…

  73. Hugging Face Daily Papers TIER_1 English(EN) ·

    Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

    Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this…

  74. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Natasha Jaques ·

    From Solo to Social Learning: Characterizing Recursive Social Improvement in LLMs

    Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent framewor…

  75. Hugging Face Daily Papers TIER_1 English(EN) ·

    Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

    As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large langu…

  76. Hugging Face Daily Papers TIER_1 English(EN) ·

    Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

    On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solut…

  77. Hugging Face Daily Papers TIER_1 English(EN) ·

    RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications

    We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning t…

  78. Hugging Face Daily Papers TIER_1 English(EN) ·

    Principled Thoughts for Latent Recursive LLM Systems

    Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. T…

  79. Hugging Face Daily Papers TIER_1 English(EN) ·

    An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

    We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-ins…

  80. arXiv cs.CV TIER_1 English(EN) · Zhixi Zhu, Kristina Gligoric ·

    Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics

    arXiv:2610.00809v1 Announce Type: new Abstract: Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on …

  81. Medium — fine-tuning tag TIER_1 English(EN) · Joya Parveen ·

    Fine-Tuning LLMs: How Do You Actually Teach an AI Model a New Skill?

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@joyaparveen0329/fine-tuning-llms-how-do-you-actually-teach-an-ai-model-a-new-skill-2645b6b6c80c?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2492/1*9t33zj8VC2fx…

  82. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    Beyond First-Order Reasoning: Why LLM Strategies Fail Under Pressure

    <p>Most AI-generated strategies suffer from the same terminal flaw: they are forward-only. If you ask an LLM to propose a market entry plan or a scaling architecture, it will present a linear progression of successful steps. It describes the ascent, but it never maps the cliffs.<…

  83. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    Moving Beyond LLM Hallucinations in Technical Analysis via Deterministic MCP Tools

    <p>Large Language Models are notoriously bad at arithmetic. When you ask a model to interpret complex financial oscillators, it isn't performing calculus; it is predicting the most likely next token based on training data. In technical analysis—where a decimal error in a volatili…

  84. Medium — fine-tuning tag TIER_1 English(EN) · SINAPSA Infocomplex ·

    LLM fine-tuning: example of a valid dataset and a chat template incorporating a Constitution for…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@contact_30070/llm-fine-tuning-example-of-a-valid-dataset-and-a-chat-template-incorporating-a-constitution-for-476b7d2f1518?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.…

  85. Medium — fine-tuning tag TIER_1 English(EN) · Harthik Mallichetty ·

    From Code Review to ARC-AGI-2: Reusing the Same ML Muscle on a Harder Reasoning Problem

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@hrmallichetty/from-code-review-to-arc-agi-2-reusing-the-same-ml-muscle-on-a-harder-reasoning-problem-a4be66501d3b?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2…

  86. Medium — fine-tuning tag TIER_1 English(EN) · Michal Molka ·

    Fine-Tuning an LLM: Teaching a Data Warehouse Schema Without RAG

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://michalmolka.medium.com/fine-tuning-an-llm-teaching-a-data-warehouse-schema-without-rag-1122ff567c41?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*Owot1Ud8ISf4rerCFiU5…

  87. dev.to — LLM tag TIER_1 English(EN) · Farhad Rahimi Klie ·

    How LLMs Work: A Deep Dive From Tokens to Intelligence

    <p>Large Language Models (LLMs) have changed how we interact with computers.</p> <p>Tools such as ChatGPT, Claude, Gemini, and many open-source models can write code, explain complex topics, summarize documents, translate languages, and generate natural-sounding conversations.</p…

  88. dev.to — LLM tag TIER_1 English(EN) · RESK ·

    LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers

    <h1> LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers </h1> <p><strong>TL;DR</strong> — LLM evaluation only produces comparable numbers when every model faces the same prompts, the same fixed judge, and the same per-axis rubrics, with a public verbatim tr…

  89. dev.to — LLM tag TIER_1 English(EN) · amrit ·

    Le Chonk Unleashed: Why Mistral’s 1‑Trillion‑Parameter Model Is Shattering LLM Benchmarks

    <h2> <strong>Mistral Large 4 (“Le Chonk”) Takes the Stage: What the Benchmarks Really Reveal</strong> </h2> <h3> A Sharper Lead </h3> <p>On <strong>October 3 2026</strong>, Mistral released a <strong>1‑trillion‑parameter</strong> model that anyone can download under an open‑weigh…

  90. dev.to — LLM tag TIER_1 English(EN) · Denis Lavrentyev ·

    Overcoming LLM Knowledge Gap: Balancing AI Tool Use with Advanced Local Model Understanding

    <h2> Introduction: The Rise of LLMs and Their Impact </h2> <p>The integration of <strong>Large Language Models (LLMs)</strong> into professional workflows has sparked a critical debate: <em>Is advanced knowledge of local LLMs essential for leveraging AI tools effectively?</em> Th…

  91. dev.to — LLM tag TIER_1 English(EN) · Filipe Martins ·

    How LLMs Work: A Journey Through One Sentence

    <h1> How LLMs Work: A Journey Through Tokens, Attention, and Transformers </h1> <p>“My favourite rock band is…”</p> <p>How does a large language model take that unfinished sentence and decide what comes next?</p> <p>You’ve probably heard that LLMs “predict the next token”. But th…

  92. dev.to — LLM tag TIER_1 English(EN) · RESK ·

    How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

    <h2> TL;DR </h2> <p>LLM evaluation only becomes useful when every model faces the same prompts, the same fixed judge, per-axis rubrics, and a public verbatim trail. The LFORLA Reverse Engineering benchmark does exactly that: it restores C source from stripped binaries, scores wit…

  93. dev.to — LLM tag TIER_1 English(EN) · Lightning Developer ·

    Scaling Intelligence: Running LLMs Across a Seven-Board ESP32-S3 Cluster

    <p>Running large language models (LLMs) on microcontrollers has long been considered a "what if" scenario reserved for theoretical discussions. However, the recent emergence of the <a href="https://github.com/Low-Zi-Hong/ESP32s3-LLM-Cluster" rel="noopener noreferrer">ESP32s3-LLM-…

  94. dev.to — LLM tag TIER_1 English(EN) · Ayi NEDJIMI ·

    LLM Fine-tuning with LoRA: A Practical Guide for Developers

    <p>Fine-tuning a large language model on your own data used to require serious GPU budgets and weeks of infrastructure work. LoRA (Low-Rank Adaptation) changed the math: you can adapt a 7B-parameter model to a specific domain in a few hours on a single consumer GPU, without touch…

  95. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    A mechanism explainer on LLM evaluation using the Team Recruitment Oracle benchmark: same prompts, a fixed judge, per-axis rubrics, and a public verbatim trail.

    A mechanism explainer on LLM evaluation using the Team Recruitment Oracle benchmark: same prompts, a fixed judge, per-axis rubrics, and a public verbatim trail. Real scores for Nemotron 3 Ultra, HY3, and our own GLM 5.2. # llm # evaluation # benchmark # ai # software # coding # d…

  96. r/ClaudeAI TIER_2 English(EN) · /u/SmirkingMan ·

    Difficult reasoning competition for LLMs - Hi-ToM

    <!-- SC_OFF --><div class="md"><p>Source: <a href="https://github.com/ying-hui-he/Hi-ToM_dataset">https://github.com/ying-hui-he/Hi-ToM_dataset</a></p> <p>I asked several LLMs to solve this Hi-ToM puzzle:</p> <p>The following story happens in chronological order. You will be give…