PulseAugur
实时 08:57:55
English(EN) Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families

AI拒绝机制仅读取模型知识的一小部分

一项新的研究论文探讨了AI模型如何处理有害请求,发现预训练后应用的对齐技术会产生浅层拒绝。研究表明,模型理解道德的能力源于其预训练阶段,形成一个独立的子空间。对齐方法随后旋转这个子空间而不是重建它,从而创建一个独立于模型更广泛道德判断的“拒绝门”。这表明当前的拒绝机制仅处理模型知识的一小部分,而其大部分理解则保持不变,易于编辑。 AI

影响 表明当前的AI安全措施可能很肤浅,可能导致更容易的越狱,并需要新的对齐策略。

排序理由 学术论文,详细介绍了AI模型行为的新发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

AI拒绝机制仅读取模型知识的一小部分

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了AI模型行为的新发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Orion Reblitz-Richardson ·

    拒绝仅读取模型知识的一小部分:跨模型家族的有害信息路由及其例外情况

    arXiv:2609.14759v1 Announce Type: cross Abstract: Alignment applied after pretraining is shallow in a measurable way: a single direction in a model's residual stream can be edited out, and the model stops refusing harmful requests. That fact says how easily refusal can be removed…