PulseAugur
EN
LIVE 08:23:41

Researchers pinpoint cross-lingual refusal circuit in multilingual MoE model

Researchers have identified a specific circuit within a multilingual Mixture-of-Experts (MoE) model, named sarvam, that is responsible for refusing harmful requests. This circuit's ability to refuse is language-invariant, meaning the underlying mechanism for detecting harm is consistent across languages. However, the actual generation of the refusal response is a late-stage process that is assembled during text generation rather than being a direct read-off from the harm detection. AI

IMPACT Identifies a specific, localizable circuit for safety refusal in multilingual MoE models, offering potential avenues for targeted repairs and cost-effective interventions.

RANK_REASON The cluster contains an academic paper detailing research findings on AI safety alignment. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Researchers pinpoint cross-lingual refusal circuit in multilingual MoE model

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Ramakrishna P. Kompella, Aadit Mahajan ·

    Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

    arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indi…