PulseAugur
实时 11:40:34
English(EN) From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

新研究优化跨不同GPU和硬件的LLM推理

研究人员正在开发新的方法来优化跨不同硬件的大型语言模型(LLM)推理和训练。Meganeura旨在通过Vulkan和Metal实现便携式GPU训练和推理,并展示出与供应商特定解决方案相比具有竞争力的性能。Celty通过共同设计GPU内核和SIMT微架构以实现稀疏矩阵-稀疏向量工作负载,专注于高效的双稀疏LLM推理。此外,一种新的分析方法无需直接测量即可估算GPU上LLM推理的能耗,有助于可持续性分析。 AI

影响 新技术有望在更广泛的硬件上实现更高效、更易于访问的LLM部署。

排序理由 多篇研究论文详细介绍了在GPU上优化LLM推理和训练的新方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

新研究优化跨不同GPU和硬件的LLM推理

报道来源 [9]

  1. arXiv cs.LG TIER_1 English(EN) · Ruokai Yin, Priyadarshini Panda ·

    Celty:用于高效双稀疏大模型推理的 SpMspV GPU 内核与 SIMT 协同设计

    arXiv:2608.01536v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly rely on sparsity to reduce inference cost, but most prior work targets a single sparsity source-either weight or activation-and optimizes for batched multi-user inference. Dual-sparsity, w…

  2. arXiv cs.LG TIER_1 English(EN) · Dzmitry Malyshau ·

    Meganeura:通过 Vulkan 和 Metal 实现便携式 GPU 训练和推理

    arXiv:2608.01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. Meganeura asks whether one compact native compiler can span both phases on consumer GPUs. Its typed static graph, automatic di…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    从 Token 到瓦时:现代 GPU 上 LLM 推理的分析性能耗估算

    The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often requires hardware telemetry, power instrumentati…

  4. Hacker News — AI stories ≥50 points TIER_1 English(EN) · Anon84 ·

    AirLLM 70B 在单块 4GB GPU 上进行推理

  5. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    国内加速卡长上下文推理性能评估

    <p>Long-context inference is becoming a core deployment scenario for large language models, and the storage and memory access bottleneck of the KV Cache directly determines the throughput and time-to-first-token (TTFT) of inference systems. The Mingxin FX100, as a domestic storag…

  6. r/LocalLLaMA TIER_1 English(EN) · /u/HomoAgens1 ·

    多个GPU在LLM推理中的扩展性如何?(试图理解基础知识)

    <!-- SC_OFF --><div class="md"><p>Hi everyone,<br /> I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs.</p> <p>Suppose I have a model running on a single GPU and then move to two or more GPUs using l…

  7. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Big Walk 就像……《无名鹅戏》很难超越。那是一次愚蠢的体验,捕捉了我所想象的成为一只感性鹅的感觉

    Big Walk is like… Untitled Goose Game is a tough act to follow. It was a silly experience that captured what I imagine it would feel like to be a sentient goose: a lot of waddling, a lot of honking, and a lot of shenanigans. That's why B… https://www. theverge.com/games/973166/bi…

  8. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    AirLLM 70B 在单块 4GB GPU 上进行推理 文章链接: https:// github.com/lyogavin/airllm 评论链接: https:// news.ycombinator.com/item?id=4 9154228 点数:

    AirLLM 70B inference with single 4GB GPU Article URL: https:// github.com/lyogavin/airllm Comments URL: https:// news.ycombinator.com/item?id=4 9154228 Points: 6 # Comments: 1 https:// github.com/lyogavin/airllm # Tech # Technology # TechNews # AI # Gadgets # Software # Cybersecu…

  9. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    llm-d 跨三大 GPU 厂商实现异构推理服务 # llmd # ai https:// twp.ai/E5EnOT

    Heterogeneous inference serving across three GPU vendors with llm-d # llmd # ai https:// twp.ai/E5EnOT