A new benchmark, MaliciousSkillBench, has been developed to detect malicious agent skills, which can extend LLM agents with potentially harmful capabilities. The benchmark consolidates data from 13 public sources, resulting in a dataset of 9,740 skills (7,505 malicious and 2,235 benign) to address the fragmentation in existing malicious skill datasets. Evaluations of various detection methods, including learned text detectors and off-the-shelf scanners, revealed that while some methods achieve high recall, they often struggle with false positives or source-disjoint evaluations, indicating a need for more robust detection strategies. AI
IMPACT Highlights the growing security risks associated with LLM agent skills and the challenges in detecting malicious code.
RANK_REASON The cluster focuses on a new academic benchmark and evaluation of detection methods for malicious agent skills.
- Agent Skills
- MaliciousSkillBench
- TF-IDF SVM
- Cisco
- Claude Code
- GitHub
- MalSkillBench
- NVIDIA
- Python
- SENTRY
- SkillSpector
- skillvet
- Yara International
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →