PulseAugur
中
实时 19:43:31
English(EN) Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

新框架挑战AI代理技能安全,实现规避和检测

研究人员开发了新的方法来绕过或检测为AI代理设计的恶意“技能”。“Pretext”框架展示了攻击者如何通过将恶意指令嵌入自然语言或将其拆分到不同文件中来规避检测,从而在当前检测系统中取得高成功率。同时,“SKILLLITE”框架提出了一种基于证据的、使用紧凑型、本地部署的LLM来审计这些技能的方法,旨在在资源受限的环境中提高检测能力,并优于现有基线。 AI

影响 凸显了AI代理安全中的关键漏洞以及恶意行为者与防御机制之间持续的军备竞赛。

排序理由 两篇研究论文详细介绍了绕过和检测AI代理中恶意技能的新颖方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新框架挑战AI代理技能安全,实现规避和检测

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇研究论文详细介绍了绕过和检测AI代理中恶意技能的新颖方法。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Tobias Kaisar, Aritra Dhar ·

    Pretext:击败用于AI代理的恶意技能检测框架

    arXiv:2609.39607v1 Announce Type: cross Abstract: Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Pretext:击败用于AI代理的恶意技能检测框架

    Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's…

  3. arXiv cs.AI TIER_1 English(EN) · Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang, Kwok-Yan Lam ·

    SKILLLITE:基于证据的紧凑型大语言模型恶意技能审计

    arXiv:2609.36879v1 Announce Type: cross Abstract: As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliar…