ENTITY
On-Policy Delta Distillation
On-Policy Delta Distillation
PulseAugur coverage of On-Policy Delta Distillation — every cluster mentioning On-Policy Delta Distillation across labs, papers, and developer communities, ranked by signal.
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D
1 day(s) with sentiment data
RECENT · PAGE 1/1 · 2 TOTAL
-
New On-Policy Delta Distillation method enhances LLM reasoning capabilities
Researchers have introduced a novel method called On-Policy Delta Distillation (OPD^2) to improve the transfer of reasoning capabilities in large language models. This technique utilizes a "delta signal," which represen…
-
New distillation method enhances LLM reasoning capabilities
Researchers have introduced On-Policy Delta Distillation (OPD²), a novel post-training method for reinforcement learning that utilizes a "delta signal" to transfer reasoning capabilities from a teacher model to a studen…