PulseAugur
实时 09:02:19
English(EN) SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

新基准测试显示LLM在从恶意软件中提取威胁情报方面存在困难

研究人员推出了SCRIPTIOC-BENCH,这是一个旨在评估大型语言模型(LLM)从基于脚本的恶意软件中提取可操作威胁情报能力的新基准测试。该基准测试包含634个手动验证的JavaScript、PowerShell和VBScript样本,对妥协指标(IOCs)进行分类,如URL、域名、IP地址和文件系统伪影。初步评估显示,即使是最先进的LLM在静态IOC恢复方面也存在困难,最高F1分数仅为65.4%。该研究还提出了一个误报分类法,以更好地理解模型错误,并探索了确定性字符串实用程序和特定任务适应等缓解措施,这些措施在恢复和精度方面显示出互补的收益。 AI

影响 强调了当前LLM在自动化恶意软件分析和威胁情报提取方面的能力局限性。

排序理由 该集群描述了一个新的学术基准测试,以及在AI安全和安全领域特定任务上对LLM的评估。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准测试显示LLM在从恶意软件中提取威胁情报方面存在困难

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个新的学术基准测试,以及在AI安全和安全领域特定任务上对LLM的评估。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Hanna Kim, Jian Cui, Minkyoo Song, Hwanjo Heo, Seungwon Shin, Kimin Lee, Xiaojing Liao ·

    SCRIPTIOC-BENCH:使用大型语言模型识别基于脚本的恶意软件中的可操作威胁情报的基准测试

    arXiv:2609.06149v1 Announce Type: cross Abstract: Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence. However, statically recovering such indicators is challenging, as …