Researchers have developed a new attack called the Groundhog Bit-Flip Attack (GBFA) that targets Mixture-of-Experts (MoE) large language models. This attack exploits the routing mechanism in MoE architectures, where specific experts can become correlated with certain tokens. By manipulating routing-layer bits, GBFA can cause significant inflation in the model's output length, extending decoding token usage by an average of 5912% across conversational, reasoning, and agentic tasks. The attack requires deactivating only a small number of experts, highlighting a robustness vulnerability in MoE designs. AI
IMPACT Reveals a significant vulnerability in MoE architectures, potentially impacting the reliability and security of large language models.
RANK_REASON Research paper detailing a new attack method against LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- agentic tasks
- conversational tasks
- experts
- Groundhog Bit-Flip Attack
- LLMs
- Mixture-of-Experts
- reasoning tasks
- routing-layer bits
- tokens
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →