A post on the Alignment Forum and LessWrong explores technical alignment for future "brain-like AGI" by drawing inspiration from human social and moral drives. The author suggests that if humans can achieve a good future, then sufficiently human-like AGIs could too, provided they possess prosocial motivations. The post delves into specific human instincts, potential failure modes like incorrect moral circles or power dynamics, and implementation details for integrating these drives into AGI code. AI
IMPACT Proposes novel approaches to AGI alignment by leveraging human social instincts, potentially guiding future research in motivation systems.
RANK_REASON The cluster consists of a detailed technical paper discussing AI alignment strategies.
- Act-based approval-directed agents
- AGI
- corrigibility
- Empowerment
- Intro to Brain-Like-AGI Safety
- Reward Function Design
- Brain-Like-AGI Safety
- Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)
- We need a field of Reward Function Design
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →