Multi-Teacher On-Policy Distillation
PulseAugur coverage of Multi-Teacher On-Policy Distillation — every cluster mentioning Multi-Teacher On-Policy Distillation across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Motif Technologies unveils 314B parameter Motif 3 LLM
Motif Technologies has released Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. The model features a novel Grouped Differential Latent At…
-
New method Soft Clamp combats AI agent over-calling of tools
Researchers have identified a failure mode in multi-teacher on-policy distillation for AI agents that use tools. This method, while improving tool-call recall, can cause agents to over-call tools inappropriately. The pa…
-
Xiaomi's MiMo-V2-Flash leads open-source coding benchmarks with efficient MoE architecture
Xiaomi has developed MiMo-V2-Flash, a 309-billion-parameter Mixture-of-Experts model that leads open-source options on SWE-Bench for coding tasks. This model achieves high performance with significantly less computation…
-
Nemotron 3 Ultra: Open-Source LLM Boasts 1M Context, 6x Throughput
Researchers have introduced Nemotron 3 Ultra, a 550 billion parameter language model that utilizes a hybrid Mamba-Transformer architecture with a Mixture-of-Experts approach. The model was trained on 20 trillion tokens …
-
New distillation method recovers LLM general capabilities after domain specialization
Researchers have developed a new method called Counteraction-Aware Multi-Teacher On-Policy Distillation (CaMOPD) to address the challenge of recovering general capabilities in large language models (LLMs) after domain s…