Researchers have identified a new supply-chain vulnerability in pretrained world models, which are used as simulators for control tasks. An adversary can embed a backdoor in a model checkpoint, allowing them to hijack downstream controllers. This attack works by subtly reshaping the latent dynamics of the model when a trigger is present, causing the controller to adopt the attacker's desired action without explicit trigger-to-action rules. The poisoned models can still pass standard clean-data diagnostics, retaining significant performance on clean tasks, and the attack's effect is temporally gated, disappearing when the trigger is removed. AI
IMPACT Highlights a new attack surface in AI control systems, potentially impacting the security of AI-driven applications.
RANK_REASON Academic paper detailing a novel attack vector on AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Dreamer
- Hugging Face
- model predictive control
- When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →