Researchers have developed a new method to attack black-box safe reinforcement learning (Safe RL) controllers used in robotics. This demonstration-guided observation attack can identify safety violations by recovering a surrogate policy and learning dynamics from demonstration transitions, without needing access to the victim network's parameters or gradients. The attack proved more effective than existing methods across various robotic tasks and budget conditions, highlighting that demonstrations used for safe learning can also serve as an attack surface. While state-adversarial regularization showed promise in defense, other tested methods provided inconsistent protection. AI
IMPACT This research highlights a potential vulnerability in deployed Safe RL systems, necessitating advancements in adversarial robustness and defense mechanisms for robotic control.
RANK_REASON This is a research paper detailing a novel attack method on Safe RL controllers. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Bullet tasks
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Jialiang Fan
- MetaDrive
- reinforcement learning
- Safe RL
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →