PulseAugur
EN
LIVE 21:36:11
ENTITY mathematical reasoning

mathematical reasoning

PulseAugur coverage of mathematical reasoning — every cluster mentioning mathematical reasoning across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
16
16 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
16
16 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/1 · 16 TOTAL
  1. RESEARCH · CL_273228 ·

    New framework Prompt2Skill optimizes LLM skills from natural language instructions

    Researchers have developed Prompt2Skill, a novel framework designed to create and optimize skills for large language models (LLMs) using only natural language instructions. This approach addresses the high cost and limi…

  2. RESEARCH · CL_269373 ·

    LLM preference generation and calibration research shows model discordance and improved methods

    Recent research explores the reliability and calibration of Large Language Model (LLM) generated preferences. One study found that while LLMs exhibit self-coherence in preferences, significant discordance exists across …

  3. TOOL · CL_268252 ·

    New Entropy Regularization Method Improves AI Model Training for Verifiable Tasks

    Researchers have identified a misalignment between standard cross-entropy (CE) training and the objective of producing correct outputs in verifiable domains like mathematical reasoning and code generation. This issue ar…

  4. TOOL · CL_254560 ·

    New 'Learning to Coach' framework enhances LLM experiential learning

    Researchers have developed a new framework called Learning to Coach (L2C) designed to improve how language models learn from experience. L2C trains a specialized LLM-as-a-Coach to distill actionable insights from an act…

  5. RESEARCH · CL_252081 ·

    New research advances on-policy distillation for LLM training · 6 sources tracked

    Researchers are developing advanced techniques for on-policy distillation (OPD), a method used to improve large language models. New approaches like $\gamma$OPD and STRIDE aim to enhance optimization stability and effic…

  6. RESEARCH · CL_239265 ·

    RISE method enhances language model training via self-extrapolation

    Researchers have introduced RISE, a novel method for improving language model post-training through self-extrapolating policy distillation. This technique constructs a synthetic teacher from the model's own reinforcemen…

  7. TOOL · CL_228622 ·

    New SHAPE framework decodes LLM math reasoning strategies

    Researchers have developed a new framework called SHAPE to analyze the Chain-of-Thought (CoT) reasoning processes of large language models (LLMs) in mathematical tasks. SHAPE examines how models interpret problems seman…

  8. RESEARCH · CL_231360 ·

    New WHALE method jointly optimizes AI agent weights and harness code

    Researchers have developed a new method called WHALE (Weight-Harness Alternating LEarning) to jointly optimize AI agent performance by simultaneously updating model weights and the harness code that manages context and …

  9. TOOL · CL_203900 ·

    New APTER framework enhances LLM reasoning with expert-grounded rubrics

    Researchers have developed APTER, a novel framework designed to enhance the reasoning capabilities of large language models in specialized domains. APTER integrates structured domain knowledge to create adaptive, expert…

  10. TOOL · CL_196097 ·

    LLM-as-a-Judge framework boosts AI reasoning with novel reward system

    Researchers have developed a novel semi-supervised learning framework that utilizes a Large Language Model (LLM) as a judge to distill knowledge into AI models. This approach employs a continuous Chain-of-Thought (CoT) …

  11. RESEARCH · CL_175929 ·

    New RLSVR method extends LLM self-improvement to open-ended tasks · 4 sources tracked

    Researchers have developed Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a new training paradigm that extends the applicability of Reinforcement Learning with Verifiable Rewards (RLVR) to open-ended tasks…

  12. TOOL · CL_148067 ·

    New SAR method extracts compact reasoning cores from LLM updates

    Researchers have developed Subspace-Aligned Rewiring (SAR), a novel post-hoc editing method for large language models. SAR identifies and isolates the core reasoning components within reinforcement learning updates, whi…

  13. RESEARCH · CL_115242 ·

    New SMMD training method enhances numerical accuracy in LLMs

    Researchers have developed a new training objective called Smooth Maximum Mean Discrepancy (SMMD) to improve the numerical precision of large language models (LLMs). Standard cross-entropy training treats numerical toke…

  14. TOOL · CL_91401 ·

    New LLM Reinforcement Learning Strategy Enhances Exploration

    Researchers have introduced Deep Dense Exploration (DDE), a novel strategy designed to improve reinforcement learning for large language models. DDE focuses on exploring deep, recoverable states within unsuccessful traj…

  15. TOOL · CL_40802 ·

    Code does not improve LLM math reasoning; structured traces do

    A new research paper explores the impact of code on mathematical reasoning in large language models. The study found that while code improves programming abilities, it does not generally enhance mathematical reasoning a…

  16. RESEARCH · CL_20433 ·

    New self-distillation methods enhance LLM reasoning and training stability

    Two new papers explore advanced self-distillation techniques for large language models, aiming to improve reasoning and efficiency. The first paper introduces "Power Distribution Bridges," which connects sampling, self-…