PulseAugur
实时 17:28:07
English(EN) Value Generalisation 1: a Research and Deployment Program

价值泛化:一种新的AI对齐方法

研究人员提出了一种新的AI对齐方法,称为“价值泛化”,旨在创建能够可靠地将人类价值观和偏好扩展到新情况的AI。这种能力被认为是当前AI系统的一个关键缺失环节,因为AI在遇到训练数据之外的场景时常常会失败。提出的方法涉及将AI的道德概念与经验概念绑定,使其道德能够随着其能力一起成长和适应,从而可能实现本质上更值得信赖和可控的“预对齐”AI。 AI

影响 这项研究可能带来更可靠、更值得信赖的AI系统,使其能够在新颖的情况下安全运行,从而解决了AI开发中的一个关键挑战。

排序理由 该集群讨论了一种新颖的AI对齐研究概念,该概念在一系列学术/研究论坛的帖子中提出。

在 Alignment Forum 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

价值泛化:一种新的AI对齐方法

报道来源 [6]

  1. Alignment Forum TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 3: Pre-aligned AIs

    <p><span>When we get explicit strong generalisation to work (see </span><a href="https://www.lesswrong.com/posts/58zFSWp8Tmxij6ckK/value-generalisation-i-an-r-and-d-program"><span>the first post</span></a><span> on the matter and </span><a href="https://www.lesswrong.com/posts/TZ…

  2. Alignment Forum TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 2: The Missing Hole in AIs’ abilities

    <h1><span>A human superpower hidden from even ourselves</span></h1><p><span>I though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it h…

  3. Alignment Forum TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 1: a Research and Deployment Program

    <p><span>I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research has become convinced t…

  4. LessWrong (AI tag) TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 3: Pre-aligned AIs

    <p><span>When we get explicit strong generalisation to work (see </span><a href="https://www.lesswrong.com/posts/58zFSWp8Tmxij6ckK/value-generalisation-i-an-r-and-d-program"><span>the first post</span></a><span> on the matter and </span><a href="https://www.lesswrong.com/posts/TZ…

  5. LessWrong (AI tag) TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 2: The Missing Hole in AIs’ abilities

    <h1><span>A human superpower hidden from even ourselves</span></h1><p><span>I though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it h…

  6. LessWrong (AI tag) TIER_1 English(EN) · Stuart_Armstrong ·

    Value Generalisation 1: a Research and Deployment Program

    <p><span>I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research has become convinced t…