PulseAugur
实时 09:18:24
English(EN) Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

研究人员精确定位多语言MoE模型中的跨语言拒绝电路

研究人员在一个名为sarvam的多语言混合专家(MoE)模型中,识别出了一个负责拒绝有害请求的特定电路。该电路的拒绝能力是语言无关的,这意味着检测有害内容的底层机制在不同语言中是一致的。然而,实际的拒绝响应生成是一个后期过程,它是在文本生成过程中组装起来的,而不是直接从有害内容检测中读取的。 AI

影响 识别出多语言MoE模型中用于安全拒绝的特定、可定位电路,为有针对性的修复和成本效益干预提供了潜在途径。

排序理由 该集群包含一篇详细介绍AI安全对齐研究成果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究人员精确定位多语言MoE模型中的跨语言拒绝电路

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Ramakrishna P. Kompella, Aadit Mahajan ·

    决定上游,延迟写入:多语言 MoE 的跨语言拒绝电路定位与定价

    arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indi…