PulseAugur
中
实时 04:12:39
English(EN) LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load

LLM 推理效率在边缘设备和云 GPU 上得到探索

两篇新研究论文探讨了高效运行大型语言模型(LLM)的挑战。第一篇论文研究了在智能手机和专用 NPU 等边缘设备上部署 LLM 的性能权衡,强调了热限制和内存带宽限制。第二篇论文介绍了一个可扩展的框架,使用启发式算法来优化异构 GPU 云环境中 LLM 推理的资源分配,旨在满足服务水平目标的同时最大限度地降低成本。 AI

影响 这些论文为优化 LLM 的性能和成本(包括设备端和云端部署)提供了见解,这对于扩展 AI 应用至关重要。

排序理由 该集群包含两篇讨论 LLM 推理性能和资源分配的学术论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

LLM 推理效率在边缘设备和云 GPU 上得到探索

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含两篇讨论 LLM 推理性能和资源分配的学术论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
125 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Shan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li, Shuo Yang, Xinyuan Tong, Yang Wang, Zhiqiang Xie, Yuwei An, Shiyi Cao, Ke Bao, Deepak Vij, Xiaoning Ding, Yichen Wang, Qingda Lu, Zhong Wang, Gao Gao, Harry Xu, Junyi Shu, Jiarong Xing, Ying Sheng ·

    Prism:通过 GPU 内存气球化实现高成本效益的多 LLM 服务

    arXiv:2505.04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall. Analysis of production traces reveals a dynam…

  2. arXiv cs.LG TIER_1 English(EN) · Pranay Tummalapalli, Sahil Arayakandy, Ritam Pal, Kautuk Kundan ·

    边缘端大语言模型推理:移动端、NPU 和 GPU 在持续负载下的性能效率权衡

    arXiv:2603.23640v2 Announce Type: replace-cross Abstract: Deploying large language models on-device for always-on personal agents demands sustained inference from hardware tightly constrained in power, thermal envelope, and memory. We benchmark Qwen 2.5 1.5B (4-bit quantised) acr…

  3. arXiv cs.LG TIER_1 English(EN) · Jiaming Cheng, Duong Tung Nguyen ·

    面向异构GPU云中受 SLO 约束的 LLM 推理的可扩展联合资源分配

    arXiv:2604.07472v2 Announce Type: replace Abstract: Serving large language model (LLM) inference in cloud environments requires jointly optimizing model selection, GPU provisioning, parallelism configuration, and workload routing under latency, accuracy, memory, and budget constr…