DeepSeek-R1-Distill-Qwen-32B
PulseAugur coverage of DeepSeek-R1-Distill-Qwen-32B — every cluster mentioning DeepSeek-R1-Distill-Qwen-32B across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
AI reasoning models create more bias than they resolve, study finds
A new research paper explores the impact of "thinking" in reasoning language models (RLMs) on fairness, specifically in high-stakes decision-making tasks. The study found that while reasoning can resolve some existing b…
-
New HIRA system enhances document classification with human-in-the-loop retrieval
Researchers have developed HIRA, a novel system designed for document classification in regulated industries. HIRA employs a training-free, on-premises retrieval-augmented cascade that combines multiple representation t…
-
LLMs struggle to maintain internal world models for complex planning tasks
Researchers have investigated why large language models struggle with planning puzzles like the Tower of Hanoi, particularly a variant where initial and goal states are complex. By training smaller Transformers on preco…
-
New Branch-Merge distillation method creates smaller, high-accuracy LLMs
Researchers have developed a new method called Branch-Merge distillation to create smaller, high-performing large language models. This approach involves selectively distilling knowledge from a large teacher model into …