ENTITY
DeepSeek-R1-Distill-Qwen-32B
DeepSeek-R1-Distill-Qwen-32B
PulseAugur coverage of DeepSeek-R1-Distill-Qwen-32B — every cluster mentioning DeepSeek-R1-Distill-Qwen-32B across labs, papers, and developer communities, ranked by signal.
Total · 30d
1
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D
1 day(s) with sentiment data
RECENT · PAGE 1/1 · 2 TOTAL
-
LLMs struggle to maintain internal world models for complex planning tasks
Researchers have investigated why large language models struggle with planning puzzles like the Tower of Hanoi, particularly a variant where initial and goal states are complex. By training smaller Transformers on preco…
-
New Branch-Merge distillation method creates smaller, high-accuracy LLMs
Researchers have developed a new method called Branch-Merge distillation to create smaller, high-performing large language models. This approach involves selectively distilling knowledge from a large teacher model into …