PulseAugur
实时 11:04:53
English(EN) Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

Mixtral MoE 在有害提示下的路由安全性分析

研究人员分析了 Mixtral 8x7B-Instruct 模型在接收良性和有害提示时的路由行为。他们利用基于激活和基于梯度的信号来理解模型如何选择专家来处理不同类型的输入。研究发现,虽然大多数专家在良性和有害提示之间共享,但一小部分专家表现出明显的偏好。抑制这些偏好专家的干预措施减少了有害响应,表明与安全相关的路由是微妙的,并且分布在各个层中。 AI

影响 为理解 Mixture-of-Experts 模型内部工作机制提供了见解,可能为未来的安全研究和开发提供信息。

排序理由 分析模型行为的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Mixtral MoE 在有害提示下的路由安全性分析

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Md Nurul Absar Siddiky ·

    Mixtral MoE 在良性与有害提示下的面向安全性的路由分析

    arXiv:2605.24270v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models activate only a small subset of parameters for each token, making router behavior a central part of model computation. This paper studies routing behavior of Mixtral 8x7B-Instruct unde…