GSM8K
PulseAugur coverage of GSM8K — every cluster mentioning GSM8K across labs, papers, and developer communities, ranked by signal.
14 day(s) with sentiment data
-
LLM tuner PolyServe reveals bugs, boosts performance with quantization
An open-source LLM tuner called PolyServe was developed to optimize model serving configurations. Benchmarking revealed several flaws in the tuner's assumptions, including a quality gate that failed to enforce its inten…
-
New method boosts LLM math reasoning with execution verification
Researchers have developed a new method for improving the mathematical reasoning capabilities of large language models by incorporating execution-based verification and dependency-aware filtering. This approach generate…
-
GLM-4-Voice uses RL to achieve state-of-the-art spoken math reasoning
Researchers have applied reinforcement learning to the GLM-4-Voice speech model to improve its mathematical reasoning capabilities. After supervised fine-tuning on spoken question-answering data, the model showed improv…
-
New YFPO framework enhances LLM reasoning with neuron-guided rewards
Researchers have introduced YFPO (Yoked Feature Preference Optimization), a novel framework designed to enhance the reasoning capabilities of large language models. This method couples response-level preference learning…
-
LoRA adapter KV cache reuse explored for quality vs. serving cost
Researchers have investigated the trade-offs between maintaining task quality and reducing serving costs when reusing the KV cache across multiple LoRA adapters in a shared backbone model. Their experiments on a Qwen3-1…
-
LLM inference optimization research details cost-quality-latency trade-offs · 2 sources tracked
Two new research papers explore the trade-offs between inference optimization techniques for large language models (LLMs), focusing on cost, quality, and latency. The first paper, "The Inference Engineering Pareto Atlas…
-
New RL Research Reveals Critical Flaw in Reward Shaping and Filtering
A new research paper highlights a critical flaw in group-relative reinforcement learning (RL) methods, specifically concerning the 'filter metric' when used with shaped rewards. The study demonstrates that if the filter…
-
Research finds truncation samplers offer limited benefit at typical LLM temperatures
A new research paper explores the impact of temperature sampling and truncation methods on large language model performance. The study found that while truncation samplers like top-p and min-p are often associated with …
-
GPT-6 Astra cracks final FrontierMath Tier 4 math problem · 1 source tracked
GPT-6 Astra has successfully solved the final remaining problem in the FrontierMath Tier 4 benchmark, a set of research-level mathematical problems designed to challenge advanced AI models. This achievement marks a sign…
-
RetroThinker framework boosts SpeechLLM reasoning accuracy
Researchers have developed RetroThinker, a novel post-training framework designed to enhance the reasoning capabilities of speech-based large language models (SpeechLLMs). This framework enables models like Moshi to sel…
-
New Circuit Reasoning Score improves RL data selection
Researchers have developed a new method called Circuit Reasoning Score (CRS) to improve data selection for reinforcement learning with verifiable rewards (RLVR). Unlike previous methods that treat data value as intrinsi…
-
Qwen2.5 model shows correlated verifier errors in math tasks · arXiv paper
A new paper investigates the independence of verifier errors within groups of completions generated by the Qwen2.5-1.5B model. Analyzing nearly 25,000 groups of eight completions across several math datasets, the study …
-
Data Scout method improves AI pretraining corpus creation
Researchers have developed a new method called Data Scout for creating specialized pretraining corpora for AI models. Unlike traditional approaches that filter large web archives, Data Scout directs targeted crawls base…
-
EGGROLL method enhances LLM training with low-rank evolution strategies
Researchers have developed EGGROLL, a method to make evolution strategies more practical for large language models by using low-rank Gaussian products instead of dense weight perturbations. This approach, while computat…
-
NCP-ArchPreview model advances language modeling with concept prediction
Researchers have introduced NCP-ArchPreview, a novel latent-space language model that moves beyond traditional next-token prediction. This model incorporates Next Concept Prediction (NCP), enabling it to learn and predi…
-
New research highlights limitations in AI evaluation methods
A new paper explores the limitations of pass@k evaluations in machine learning, particularly when extrapolating beyond the number of samples collected. The research demonstrates that fixed-n success counts in conditiona…
-
Spark-X2.5 LLM Tops Hugging Face Charts, Launches Math Reasoning Challenge
Spark-X2.5, a large language model, has achieved the top spot on Hugging Face's trending models list. To further test its capabilities, the developers are launching the Spark-X2.5 Math Reasoning Challenge. Participants …
-
New framework evaluates open LLMs on performance, latency, and memory
A new research paper proposes a unified evaluation framework for open reasoning language models, moving beyond simple accuracy metrics. The study tested seven model configurations across four benchmarks, analyzing not o…
-
New framework boosts MoE model inference efficiency
Researchers have developed a cache-aware framework to improve the memory efficiency of Mixture-of-Experts (MoE) models during inference. The proposed post-training method jointly adapts the MoE backbone and lightweight …
-
LLM benchmark contamination inflates scores but rarely reorders leaderboards
A new paper from arXiv investigates benchmark contamination in large language models, distinguishing between score inflation and leaderboard reordering. The research found that while contamination does inflate absolute …