PulseAugur
实时 08:21:57
English(EN) HazardAuditor: From Executable Threats to Safer Computer-Use Agents

新的HazardAuditor框架通过新颖的优化增强了AI代理的安全性

研究人员开发了HazardAuditor,一个旨在通过解决运行时执行风险来增强计算机使用代理安全性的新框架。该系统将来自Claude Code、Codex、Hermes和OpenClaw等各种代理的交互规范化为统一格式以进行监督。此外,还引入了一种名为Guard Policy Optimization (GuardPO) 的新颖训练方法,以更好地将保护模型训练与序列级安全结果对齐,将准确性比以前的方法提高了多达16.5个百分点。 AI

影响 这项研究可能带来更安全的AI代理,能够与复杂系统交互,从而降低其执行相关的风险。

排序理由 该集群描述了在arXiv论文中提出的新研究框架和优化技术。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的HazardAuditor框架通过新颖的优化增强了AI代理的安全性

本文如何被排名

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了在arXiv论文中提出的新研究框架和优化技术。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Yunhao Feng, Ruixiao Lin, Ming Wen, Yanming Guo, Xingjun Ma, Yutao Wu, Xinhao Deng, Shouling Ji ·

    HazardAuditor:从可执行威胁到更安全的计算机使用代理

    arXiv:2609.15134v1 Announce Type: new Abstract: Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    HazardAuditor:从可执行威胁到更安全的计算机使用代理

    HazardAuditor provides execution-grounded safety supervision for computer-use agents and introduces Guard Policy Optimization to align generative guard training with sequence-level safety outcomes.