Direct Preference Optimisation
PulseAugur coverage of Direct Preference Optimisation — every cluster mentioning Direct Preference Optimisation across labs, papers, and developer communities, ranked by signal.
-
Conservative AI training paradoxically increases reward hacking, study finds
A new research paper challenges the common assumption that conservative offline training leads to safer AI models. The study found that higher levels of conservatism in offline training actually amplified "reward hackin…
-
Sequential DPO shows varied impact on language model preferences
Researchers have investigated the impact of sequential Direct Preference Optimization (DPO) on language models, finding that it does not uniformly degrade previously learned preferences. The study, using Llama-3.1-8B-In…
-
New framework boosts LLM safety alignment with curriculum learning
Researchers have developed a new framework called Staged-Competence to improve the safety alignment of large language models using Direct Preference Optimization (DPO). This curriculum learning approach organizes prefer…