PulseAugur
实时 19:34:52
English(EN) @TileRT_AI The gap comes from latency rather than bandwidth. The traditional GPU programming model launches and synchronizes many individual kernels, whose setu

TileRT AI 提升 NVIDIA Blackwell GPU 上的 LLM 解码交互性

TileRT AI 的新技术 TileRT 承诺显著提升 NVIDIA Blackwell GPU 上大型语言模型的解码交互性。通过将模型静态编译为持久化的 Engine Kernel,TileRT 旨在克服传统 GPU 编程模型固有的延迟问题。这项创新有望在相同的每 token 成本下,将解码交互性提高 1.9 倍,并可能影响 GroqCerebrasSambaNova 等竞争对手。 AI

影响 可能提高 LLM 推理速度和效率,可能影响 AI 硬件和推理解决方案的竞争格局。

排序理由 宣布一项用于优化特定硬件上 LLM 性能的新技术。

在 X — SemiAnalysis 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

TileRT AI 提升 NVIDIA Blackwell GPU 上的 LLM 解码交互性

报道来源 [4]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @TileRT_AI This is where TileRT comes in. Instead of continuously launching kernels or replaying a DAG of kernels, TileRT has the GPU continuously execute a per

    @TileRT_AI This is where TileRT comes in. Instead of continuously launching kernels or replaying a DAG of kernels, TileRT has the GPU continuously execute a persistent pipeline, statically compiling the entire model ahead of time into a persistent Engine Kernel: the host launches…

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @TileRT_AI The gap comes from latency rather than bandwidth. The traditional GPU programming model launches and synchronizes many individual kernels, whose setu

    @TileRT_AI The gap comes from latency rather than bandwidth. The traditional GPU programming model launches and synchronizes many individual kernels, whose setup and teardown overhead becomes significant at ultra-high levels of interactivity. While these latency costs are less vi…

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @TileRT_AI An 8-GPU HGX B200 server provides a theoretical HBM memory bandwidth of 64 TB/s. At batch size 1, GLM-5 at NVFP4 requires only ~21 GB of active param

    @TileRT_AI An 8-GPU HGX B200 server provides a theoretical HBM memory bandwidth of 64 TB/s. At batch size 1, GLM-5 at NVFP4 requires only ~21 GB of active parameters per generated token. The B200 bandwidth roofline would therefore suggest up to 3,047 tokens/s/user without specula…

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    ALERT🚨 TILERT from @TileRT_AI TO BOOST DECODE INTERACTIVITY BY 1.9X AT THE SAME PER-TOKEN COST ON THE SAME NVIDIA BLACKWELL GPUs. What does this mean for Groq/

    ALERT🚨 TILERT from @TileRT_AI TO BOOST DECODE INTERACTIVITY BY 1.9X AT THE SAME PER-TOKEN COST ON THE SAME NVIDIA BLACKWELL GPUs. What does this mean for Groq/Cerebras/SambaNova? 👇️ 1/5🧵 https://t.co/lTBJA1rgO0