Qwen2.5-Math-1.5B
PulseAugur coverage of Qwen2.5-Math-1.5B — every cluster mentioning Qwen2.5-Math-1.5B across labs, papers, and developer communities, ranked by signal.
-
New ReCo method improves GRPO for language model reasoning
Researchers have developed ReCo, a novel reweighting method designed to improve Group Relative Policy Optimization (GRPO) in language models. GRPO, a standard reinforcement learning technique, has been observed to somet…
-
New method uses wrong drafts to boost LLM math capabilities
Researchers have developed a novel technique called "Weak-to-Strong Elicitation via Mismatched Wrong Drafts" to improve the capabilities of large language models. This method involves using mathematically incorrect draf…
-
New research advances policy optimization for robotics and LLMs
Researchers have introduced several new methods to enhance policy optimization in reinforcement learning, particularly for complex tasks involving robotics and large language models. MODIP aims to efficiently fine-tune …
-
Qwen2.5-Math-1.5B model fine-tuned for mathematical tasks
A technical guide details the process of fine-tuning the Qwen2.5-Math-1.5B model. The article outlines the steps involved in adapting this specific language model for mathematical tasks, likely to improve its performanc…
-
New RL methods boost LLM reasoning and efficiency
Two new research papers introduce novel reinforcement learning techniques for enhancing language model reasoning. The first, GAGPO, proposes a critic-free method for precise temporal credit assignment in multi-turn envi…
-
New research tackles multilingual models, efficient inference, and data contamination
Recent research explores various facets of language model development and application. Google DeepMind's ATLAS project introduces new scaling laws for multilingual models, aiming to optimize training for languages beyon…