PulseAugur
中
实时 08:17:39
English(EN) Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

新攻击目标大语言模型服务框架,绕过模型防御

研究人员开发了一种新的方法来攻击大语言模型(LLM)的服务框架,而不是模型本身。这种“填充与挤压”(Fill and Squeeze)策略通过耗尽 KV 缓存并强制重复抢占来攻击调度器的状态转换。该攻击在 vLLM 框架上显著降低了首次令牌生成时间(TTFT)和每个输出令牌的时间,在实际的黑盒环境中以比以前更低的成本证明了其有效性。 AI

影响 这项研究突显了大语言模型服务基础设施中一类新的漏洞,可能影响部署安全性和成本效益。

排序理由 详细介绍大语言模型服务框架新攻击方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新攻击目标大语言模型服务框架,绕过模型防御

本文如何被排名

Signal score
18 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍大语言模型服务框架新攻击方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tianyi Wang, Huawei Fan, Yuanchao Shu, Peng Cheng, Cong Wang ·

    重新思考延迟拒绝服务攻击:攻击大型语言模型服务框架而非模型本身

    arXiv:2602.07878v2 Announce Type: replace-cross Abstract: LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting …