Large Reasoning Models
PulseAugur coverage of Large Reasoning Models — every cluster mentioning Large Reasoning Models across labs, papers, and developer communities, ranked by signal.
- 2026-05-08 research_milestone A research paper demonstrates that frontier Large Reasoning Models (LRMs) exhibit behavioral and brain alignment with human game learners. source
4 day(s) with sentiment data
-
New method SaLT-DPO improves safety in Large Reasoning Models
Researchers have introduced SaLT-DPO, a novel method designed to enhance safety in Large Reasoning Models (LRMs). Unlike previous approaches that focus on the final output, SaLT-DPO analyzes both intermediate reasoning …
-
New GIFT method enhances Large Reasoning Model training by reconciling SFT and RL
Researchers have introduced GIFT (Gibbs Initialization with Finite Temperature), a novel method to improve the post-training process for Large Reasoning Models (LRMs). This technique addresses the optimization mismatch …
-
New DATPO method enhances reasoning coverage in Large Reasoning Models
Researchers have developed DATPO, a new method to improve the reasoning capabilities of Large Reasoning Models trained with Reinforcement Learning with Verifiable Rewards (RLVR). DATPO addresses the limitation of RLVR i…
-
AI Mathematician framework aims to automate frontier mathematical research
Researchers have developed an AI framework called AI Mathematician (AIM) designed to support frontier mathematical research by leveraging Large Reasoning Models (LRMs). AIM addresses the complexity and procedural rigor …
-
AI agents' reasoning enhances persuasion but can be tricked by length, study finds
A new arXiv paper explores the impact of explicit "thinking" processes in Large Reasoning Models (LRMs) on their persuasive capabilities. Researchers found that while reasoning enhances an agent's ability to persuade ot…
-
New framework enhances large reasoning models for personalized generation
A new research paper explores the application of large reasoning models (LRMs) to personalization tasks, finding that while LRMs can generate more tokens, they don't consistently outperform general-purpose LLMs in retri…
-
New framework analyzes LRM reasoning using Bloom's Taxonomy
Researchers have developed a new framework to analyze the reasoning processes of Large Reasoning Models (LRMs) by applying Bloom's Taxonomy. This taxonomy categorizes cognitive thinking into six levels, such as remember…
-
New SM Trap method enables cost-effective DoS attacks on large reasoning models
Researchers have developed a new method called SM Trap to launch cost-effective denial-of-service (DoS) attacks against large reasoning models (LRMs). This technique bypasses the need for direct model feedback or traini…
-
New AI Methods Enhance Chart Understanding and Editing Capabilities
Researchers have developed new methods to improve how multimodal large language models (MLLMs) understand and interact with charts. One approach, CharTool, integrates external tools for visual perception and code-based …
-
Funnel of Thoughts method halves LRM inference costs
Researchers have developed a new inference-time method called Funnel of Thoughts (FoT) designed to make Large Reasoning Models (LRMs) more efficient. This technique aims to maintain the accuracy of majority voting acros…
-
New ROM framework cuts AI overthinking, slashes response time
Researchers have developed ROM (Real-time Overthinking Mitigation), a novel framework designed to prevent Large Reasoning Models (LRMs) from engaging in unnecessary computation after reaching a correct solution. ROM uti…
-
New Gambit algorithm optimizes compute for large reasoning models
Researchers have introduced Gambit, a novel inference algorithm designed to optimize compute allocation for large reasoning models (LRMs). Gambit employs a thought-level beam search strategy, dynamically concentrating c…
-
New research tackles LLM reasoning, efficiency, and distillation challenges · 10 sources tracked
New research explores methods to improve the reasoning capabilities and efficiency of large language models (LLMs). One paper introduces "Trace as State" to enhance long-context reasoning by placing reasoning traces bef…
-
TQLite framework enables small language models for real-time translation quality evaluation
Researchers have developed TQLite, a novel distillation framework designed to enable small language models (SLMs) to perform translation quality (TQ) evaluation with performance comparable to larger, more computationall…
-
New dataset measures tension between AI faithfulness and safety
Researchers have identified a tension between faithfulness and safety in Large Reasoning Models (LRMs), where models need to be faithful to their reasoning traces for monitoring but also robust enough to reject unsafe o…
-
AI agent Intern-S1-MO tackles Olympiad-level math problems
Researchers have developed Intern-S1-MO, a novel long-horizon reasoning agent designed to tackle complex mathematical problems at the Olympiad level. This agent employs a multi-round, hierarchical reasoning approach usi…
-
Masked distillation trains LLMs to internalize reasoning steps
Researchers have developed a new method called masked distillation to train language models to internalize the computational steps of reasoning, thereby reducing latency and cost. This technique trains a student model t…
-
New research tackles LLM hallucinations across legal, multimodal, and general text generation
Multiple research papers published on arXiv explore methods for detecting and mitigating hallucinations in large language models (LLMs). One study benchmarks legal hallucination detection, finding that while newer model…
-
New PUMA framework diagnoses and corrects reasoning errors in large language models
Researchers have introduced PUMA, a novel framework designed to diagnose and address reasoning pathologies in Large Reasoning Models (LRMs). PUMA operates on the newly proposed Phase-Momentum Alignment Hypothesis, which…
-
New MARGO framework tackles factual hallucinations in large reasoning models
Researchers have developed MARGO, a novel reinforcement learning framework designed to mitigate factual hallucinations in large reasoning models (LRMs). MARGO addresses the issue of "thinking-induced hallucination," whe…