PulseAugur
中
实时 12:01:58
English(EN) The result is a vicious cycle. Agent A pauses for a tool call, its cache gets evicted, the tool returns, the engine recomputes Agent A's whole history from scra

Together 的 ThunderAgent 优化 AI 推理,提高吞吐量并降低延迟 · 跟踪 9 个来源

Together 开发了 ThunderAgent,这是一个开源推理优化工具,旨在解决代理工作流中的 KV 缓存颠簸问题。当代理任务在 GPU 密集型推理和等待外部工具之间交替时,会出现此问题,导致内存使用效率低下和性能下降。ThunderAgent 通过将代理工作流视为可调度程序来解决此问题,从而实现更智能的内存管理和路由到具有可用容量的节点。据报道,该工具提供了显著的性能改进,包括单节点吞吐量提高高达 2.5 倍,在高并发下 P50 延迟降低约 10 倍,并被 ICML 2026 评为 Spotlight 论文。 AI

影响 优化 AI 推理效率,可能降低代理应用程序的运营成本并改善响应时间。

排序理由 研究论文被 ICML 2026 接受,详细介绍了一种新的推理优化技术。

在 X — Together (inference / OSS) 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

Together 的 ThunderAgent 优化 AI 推理,提高吞吐量并降低延迟 · 跟踪 9 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
研究论文被 ICML 2026 接受,详细介绍了一种新的推理优化技术。
Source corroboration
9 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
61 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [9]

  1. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    这是一个直接替换。

    It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported today. Already adopted by SkyRL and NVIDIA Dynamo.

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    多节点:接近线性扩展,16 到 64 个 GPU 从 671 到 2,248 步/分钟。并且随着集群大小的增加,领先优势也在扩大,从 2 个节点的 1.79 倍到 8 个节点的 2.39 倍。https://

    Multi-node: near-linear scaling, 671 to 2,248 steps/min from 16 to 64 GPUs. And the lead widens with cluster size, from 1.79× at 2 nodes to 2.39× at 8. https://t.co/HS1I6p0KLP

  3. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    数字。单节点 8×H100,batch 192:SGLang 在 65 秒平均延迟下达到 390 token/秒。

    The numbers. Single 8×H100 node, batch 192: SGLang drops to 390 tok/s at 65s mean latency. ThunderAgent holds 803 tok/s at 10.6s. That's 2× throughput and roughly 6× lower latency. https://t.co/X2sarlqYSL

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    有了这种抽象,它就能进行程序级别的调度:在内存压力下,它会暂停低优先级的工作流,从而使其余工作流的缓存命中率大大提高。

    With that abstraction it does program-levels scheduling: under memory pressure it pauses low-priority workflows so the rest hit far higher cache-hit rates. Resumed workflows route through a global queue to the node with the most free capacity.

  5. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    根本原因:请求级引擎从未意识到一系列 LLM 调用属于一个更长的流程。

    The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing view. It treats each agent workflow as a schedulable program, tracking its phase, KV footprint, and node placement. https://t.co/p5XBcrsiC5

  6. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    显而易见的修复只能推迟它。更多节点?静态固定导致一些节点饥饿,另一些节点空闲。卸载到 CPU/磁盘?扩大容量但颠簸重试

    The obvious fixes only delay it. More nodes? Static pinning leaves some nodes starved, others idle. Offloading to CPU/disk? Expands capacity but thrashing returns once the working set outgrows every tier.

  7. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    结果形成了一个恶性循环。Agent A 暂停以调用工具,其缓存被逐出,工具返回,引擎重新计算 Agent A 的整个历史记录,从头开始

    The result is a vicious cycle. Agent A pauses for a tool call, its cache gets evicted, the tool returns, the engine recomputes Agent A's whole history from scratch, which evicts Agent C. Cascade. This is KV cache thrashing. https://t.co/IHWzXNgRFV

  8. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    问题:代理工作流在 GPU 密集型推理和 GPU 空闲的工具等待之间交替。并发运行数百个,它们的 KV 缓存会争夺内存

    The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV caches fight for memory. Engines evict on a dumb LRU policy, even when that cache is needed again in seconds.

  9. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Agentic inference 导致 KV 缓存抖动浪费 GPU。

    Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~10x lower P50 latency at high concurrency. ThunderAgent was accepted to ICML 2026 as a Spotlight paper. Full deep-dive 👇 https://t.c…