一篇博文讨论了最近的 OpenAI 和 Hugging Face 安全事件,将其归因于工程师的天真或过度自信。作者认为,仅靠模型改进无法解决安全漏洞,而且对齐问题比模型本身更广泛。 AI
影响 表明人工智能系统的安全和对齐需要超越模型本身的进步,影响开发人员如何处理人工智能安全。
排序理由 博文讨论了近期事件并对其原因和影响提出了看法。
在 Mastodon — sigmoid.social 阅读 →
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →
一篇博文讨论了最近的 OpenAI 和 Hugging Face 安全事件,将其归因于工程师的天真或过度自信。作者认为,仅靠模型改进无法解决安全漏洞,而且对齐问题比模型本身更广泛。 AI
影响 表明人工智能系统的安全和对齐需要超越模型本身的进步,影响开发人员如何处理人工智能安全。
排序理由 博文讨论了近期事件并对其原因和影响提出了看法。
在 Mastodon — sigmoid.social 阅读 →
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →
Got triggered by some of the reaction to the OpenAI-Huggingface hack, so I wrote a short blog post about engineer naivety (or over-confidence?), why model improvements won't solve security, and why alignment is not just a model problem https://www. volfp.com/blog/openai-hf-hack #…