PulseAugur
EN
LIVE 11:04:41
ENTITY Direct On-Policy Distillation

Direct On-Policy Distillation

PulseAugur coverage of Direct On-Policy Distillation — every cluster mentioning Direct On-Policy Distillation across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
  1. RESEARCH · CL_139531 ·

    New framework enhances LLM training by reducing noise in weaker models

    Researchers have developed a new framework called Contrastive Weak-to-Strong Generalization (ConG) to improve the training of large language models. ConG addresses limitations in existing weak-to-strong generalization m…

  2. RESEARCH · CL_128417 ·

    New research explores controllable generalization failures and efficient RL distillation for LLMs

    Researchers are exploring new methods to improve language model generalization and reasoning capabilities. One paper proposes a technique to construct models that exhibit controllable generalization failures by training…