A new paper proposes reformulating the AI alignment problem as a social choice issue, moving beyond standard reinforcement learning from human feedback. The research suggests that by focusing on an algorithm's welfare consequences, alignment can be approached using linear optimization and tools from welfare economics and mechanism design. This framework allows for translating alignment protocols into welfare outcomes and vice versa, with empirical demonstrations using human preferences on various scenarios like kidney allocation and LLM responses. AI
影响 Proposes a novel theoretical framework for AI alignment that could lead to more robust and ethically sound AI systems.
排序理由 Academic paper published on arXiv detailing a new theoretical approach to AI alignment. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- charitable food distribution
- mechanism design
- random-dictatorship
- reinforcement learning from human feedback
- trolley problems
- voting-by-issues
- welfare economics
- Zachary Wojtowicz
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →