PulseAugur
EN
LIVE 17:27:42

Value Generalisation: A New Approach to AI Alignment

Researchers are proposing a new approach to AI alignment called "value generalisation," which aims to create AIs that can reliably extend human values and preferences to novel situations. This capability is seen as a critical missing piece in current AI systems, which often fail when encountering scenarios outside their training data. The proposed method involves binding an AI's moral concepts to empirical concepts, allowing its morality to grow and adapt alongside its capabilities, potentially enabling "pre-aligned" AIs that are inherently more trustworthy and controllable. AI

IMPACT This research could lead to more reliable and trustworthy AI systems capable of operating safely in novel situations, addressing a key challenge in AI development.

RANK_REASON The cluster discusses a novel research concept for AI alignment, presented in a series of posts on academic/research forums.

Read on Alignment Forum →

AI-generated summary · Google Gemini · from 6 sources. How we write summaries →

Value Generalisation: A New Approach to AI Alignment

COVERAGE [6]

  1. Alignment Forum TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 3: Pre-aligned AIs

    <p><span>When we get explicit strong generalisation to work (see </span><a href="https://www.lesswrong.com/posts/58zFSWp8Tmxij6ckK/value-generalisation-i-an-r-and-d-program"><span>the first post</span></a><span> on the matter and </span><a href="https://www.lesswrong.com/posts/TZ…

  2. Alignment Forum TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 2: The Missing Hole in AIs’ abilities

    <h1><span>A human superpower hidden from even ourselves</span></h1><p><span>I though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it h…

  3. Alignment Forum TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 1: a Research and Deployment Program

    <p><span>I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research has become convinced t…

  4. LessWrong (AI tag) TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 3: Pre-aligned AIs

    <p><span>When we get explicit strong generalisation to work (see </span><a href="https://www.lesswrong.com/posts/58zFSWp8Tmxij6ckK/value-generalisation-i-an-r-and-d-program"><span>the first post</span></a><span> on the matter and </span><a href="https://www.lesswrong.com/posts/TZ…

  5. LessWrong (AI tag) TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 2: The Missing Hole in AIs’ abilities

    <h1><span>A human superpower hidden from even ourselves</span></h1><p><span>I though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it h…

  6. LessWrong (AI tag) TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 1: a Research and Deployment Program

    <p><span>I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research has become convinced t…