ENTITY
Direct On-Policy Distillation
Direct On-Policy Distillation
PulseAugur coverage of Direct On-Policy Distillation — every cluster mentioning Direct On-Policy Distillation across labs, papers, and developer communities, ranked by signal.
Total · 30d
0
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
New framework enhances LLM training by reducing noise in weaker models
Researchers have developed a new framework called Contrastive Weak-to-Strong Generalization (ConG) to improve the training of large language models. ConG addresses limitations in existing weak-to-strong generalization m…
-
New research explores controllable generalization failures and efficient RL distillation for LLMs
Researchers are exploring new methods to improve language model generalization and reasoning capabilities. One paper proposes a technique to construct models that exhibit controllable generalization failures by training…