Researchers have developed APTInvestBench, a new benchmark designed to evaluate the robustness of large language model (LLM) agents in investigating advanced persistent threats (APTs) across different telemetry settings. The benchmark includes 370 cases derived from 56 attack reconstructions, totaling over 16 million log records, and assesses agents' ability to gather sufficient evidence and provide formal citations. Initial tests with eleven LLMs showed that agents could acquire sufficient evidence for an average of 44.3% of recoverable attack actions, but only 25.0% were supported by formal citations, highlighting a significant gap in reliable evidence acquisition and reporting. AI
IMPACT This benchmark could accelerate the development of more reliable AI agents for cybersecurity, improving threat detection and response capabilities.
RANK_REASON The item is a research paper detailing a new benchmark for evaluating AI agents in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →