Qwen3-Base
PulseAugur coverage of Qwen3-Base — every cluster mentioning Qwen3-Base across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New ShamAN-Q method slashes LLM weights to sub-1-bit
Researchers have developed ShamAN-Q, a novel sub-1-bit post-training quantization method for large language models. This technique enhances NanoQuant by incorporating a dense curvature metric derived from the Shampoo op…
-
New SP3O method mitigates Value Flattening in PPO for LLMs
Researchers have identified a failure mode in Proximal Policy Optimization (PPO) called Value Flattening, where state values estimated by a critic become flat despite sharp changes across intermediate states. This issue…
-
New framework unifies on-policy self-distillation for LLM reasoning · 3 sources tracked
Researchers have developed a unified framework for on-policy self-distillation (OPSD) to enhance LLM reasoning by integrating privileged information into model parameters. This new framework, Unified On-Policy Self-Dist…
-
New replay method boosts GRPO training for LLMs
Researchers have developed a new method for improving the sample efficiency of GRPO, a reinforcement learning technique used for training large language models. The proposed rollout-level experience replay buffer stores…