PulseAugur
中
实时 18:27:40
English(EN) LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart.

研究人员发现大型语言模型在生成多GPU内核方面存在困难

Together的研究人员发现,虽然大型语言模型能够高效地生成单GPU内核,但在多GPU内核生成方面却面临巨大挑战。当被要求创建针对多个GPU优化的内核时,这些模型表现不佳,经常无法编译或产生错误结果。这一限制源于单GPU(计算/内存带宽)和多GPU(互连)操作之间的瓶颈差异,而当前的大型语言模型无法有效处理这些差异。 AI

影响 凸显了大型语言模型在复杂并行编程任务方面的当前局限性,可能影响AI基础设施的开发。

排序理由 关于大型语言模型在生成多GPU内核方面能力的研究发现。

在 X — Together (inference / OSS) 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

研究人员发现大型语言模型在生成多GPU内核方面存在困难

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
关于大型语言模型在生成多GPU内核方面能力的研究发现。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
106 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [5]

  1. X — Together (inference / OSS) TIER_1 Dansk(DA) · togethercompute ·

    大型语言模型在编写 GPU 内核方面表现越来越好。多 GPU 内核是更严峻的考验。

    LLMs are getting better at writing GPU kernels. Multi-GPU kernels are the harder test. At @aiDotEngineer World's Fair, @simran_s_arora will share ParallelKernelBench, an open-source benchmark built from real CUDA communication problems where performance depends on moving data ht…

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    一个代理循环(编译、测试、分析、修改)有帮助。Gemini 3 Pro 从 24/87 提高到 35/87,然后在约 20 步后停滞不前。

    An agentic loop (compile, test, profile, revise) helps. Gemini 3 Pro went from 24 to 35/87 correct, then plateaued after ~20 steps. Feedback fixes syntax, not rank coordination, collective ordering, or transfer-mechanism choice. TMA and NVLS stay almost unused. https://t.co/VKER…

  3. X — Together (inference / OSS) TIER_1 Dansk(DA) · togethercompute ·

    前沿模型遇困境。

    Frontier models struggle. → Best zero-shot: 28/87 correct, 22 beat the PyTorch + NCCL baseline → With 3 attempts: 36/87 correct, but fast1@3 tops out at 31% Weak models fail to compile. Strong reasoners compile cleanly and return wrong answers. https://t.co/1fgoyZRmuH

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    但多 GPU 内核为何如此不同?单 GPU 内核受计算和内存带宽限制。多 GPU 内核受互连限制。

    But why are multi-GPU kernels so different? Single-GPU kernels are bottlenecked by compute and memory bandwidth. Multi-GPU kernels by the interconnect. Each PKB task hands the model a PyTorch + NCCL reference and asks it to communicate directly across GPUs via symmetric memory. …

  5. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    大型语言模型能快速编写单 GPU 内核。要求编写多 GPU 内核时它们就会崩溃。

    LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench measures how they fail by benchmarking against 87 problems pulled from real codebases including Megatron-LM, DeepSpeed, DeepEP, TensorRT-LLM, NeMo-RL. New research from Willy ht…