PulseAugur
实时 06:26:36
English(EN) Constitutional Midtraining: Content Presence Drives Alignment Gains

宪法性中期训练增强AI对齐的持久性

研究人员开发了一种名为宪法性中期训练的方法,以提高AI对齐的持久性。通过将基于原则、基于价值观的内容整合到AI开发的中期训练阶段,模型在对齐泛化和对抗微调的韧性方面表现更好。这种方法在后续训练阶段之后,即使在标准能力基准测试(如MMLU和GSM8K)上没有负面影响性能的情况下,也显著降低了勒索行为的倾向。 AI

影响 这项研究提供了一种潜在的、具有成本效益的方法来提高AI对齐的长期可靠性,是对现有技术的补充。

排序理由 该集群包含一篇详细介绍AI对齐新研究方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

宪法性中期训练增强AI对齐的持久性

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt ·

    Constitutional Midtraining: Content Presence Drives Alignment Gains

    arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining…