Qwen2.5-Coder-3B
PulseAugur coverage of Qwen2.5-Coder-3B — every cluster mentioning Qwen2.5-Coder-3B across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
EvoSQL framework enhances Text-to-SQL with critic-generator co-evolution
Researchers have developed EvoSQL, a novel framework designed to enhance Text-to-SQL capabilities by treating SQL synthesis as an iterative process between a generator and a critic. This system incorporates a memory com…
-
Chinese researchers release VibeThinker-3B, a compact 3B model matching larger models
Chinese researchers have developed VibeThinker-3B, a compact 3-billion parameter dense reasoning model. This model, built upon Qwen2.5-Coder-3B and utilizing Spectrum-to-Signal training, achieves performance comparable …
-
AI research: SFT overtraining causes rank inversion in code generation models
A new research paper explores the phenomenon of supervised fine-tuning (SFT) overtraining in reinforcement learning from human feedback (RLHF) for code generation models. The study, focusing on Qwen2.5-Coder-3B and Deep…
-
New 3B model VibeThinker matches frontier math & coding performance
Researchers have developed VibeThinker-3B, a compact 3-billion parameter model that achieves performance comparable to much larger models in mathematics and coding tasks. This model, built upon Qwen2.5-Coder-3B and util…
-
New AI framework trains code models to self-correct security flaws
Researchers have developed a novel framework called Tree Self-Play (TSP) to address the inherent security vulnerabilities in large language models trained on code. Current methods like supervised fine-tuning and reinfor…