PulseAugur
中
实时 07:33:33
English(EN) APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry

新基准评估用于网络威胁调查的大语言模型代理

研究人员开发了APTInvestBench,这是一个新的基准,旨在评估大型语言模型(LLM)代理在不同遥测设置下调查高级持续性威胁(APT)的鲁棒性。该基准包含370个案例,源自56次攻击重建,总计超过1600万条日志记录,并评估代理收集足够证据和提供正式引用的能力。对11个LLM的初步测试表明,代理能够为平均44.3%的可恢复攻击行为获取足够证据,但只有25.0%的证据得到了正式引用的支持,这凸显了可靠证据获取和报告方面存在重大差距。 AI

影响 该基准可以加速开发更可靠的网络安全AI代理,提高威胁检测和响应能力。

排序理由 该项目是一篇研究论文,详细介绍了一个用于在特定领域评估AI代理的新基准。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准评估用于网络威胁调查的大语言模型代理

本文如何被排名

Signal score
21 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目是一篇研究论文,详细介绍了一个用于在特定领域评估AI代理的新基准。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yu Wang, Shuhao Li, Tao Yin, Ziyang Li, Xueying Zhao, Peishuai Sun, Jiang Xie ·

    APTInvestBench:在不同遥测数据下评估自主APT调查

    arXiv:2609.38954v1 Announce Type: cross Abstract: Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry…