DeepScaleR
PulseAugur coverage of DeepScaleR — every cluster mentioning DeepScaleR across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New RL method NGU tackles 'Matthew Effect' in LLM training · 2 sources tracked
A new research paper introduces the "Matthew Effect in RL for LLMs," observing that reinforcement learning disproportionately benefits easy tasks for large language models, while hard tasks see minimal improvement. To a…
-
New MADA-RL framework boosts compact model reasoning with parameter-efficient debate learning
Researchers have developed MADA-RL, a novel post-training framework designed to enhance the reasoning capabilities of compact language models (under 4 billion parameters) using parameter-efficient methods. This framewor…
-
New method speeds up RLHF training with adaptive parallelism
Researchers have developed a new method called PAT to accelerate the training of Reinforcement Learning from Human Feedback (RLHF) models. This technique dynamically adjusts tensor parallelism during the generation stag…