PulseAugur
中
实时 10:11:39

新研究探索通过射频广播和稀疏性技术实现高效LLM推理

研究人员正在开发新颖的方法来提高大型语言模型(LLM)推理的效率。一种方法AIR-LLM提出通过射频广播LLM权重,使边缘设备无需存储权重即可进行推理,从而显著降低能耗和通信时长。其他研究则侧重于通过稀疏性技术优化推理,例如SparseEngine和TopK-Guided,它们通过选择性地处理模型的一部分来降低内存和计算成本。此外,正在开发MINCE和evalstats等新框架,通过缩小数据集和提高LLM评判分数统计分析的可靠性来简化LLM评估。 AI

影响 这些进展旨在使LLM在边缘设备上更易于访问和更高效,并降低推理和评估的计算成本。

排序理由 多篇研究论文介绍了LLM推理优化和评估的新颖方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 25 个来源。 我们如何撰写摘要 →

新研究探索通过射频广播和稀疏性技术实现高效LLM推理

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了LLM推理优化和评估的新颖方法。
Source corroboration
25 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
12 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+11 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [25]

  1. arXiv cs.AI TIER_1 Español(ES) · Tate Berenbaum (Not Community Labs Inc.), Matias Parij (Not Community Labs Inc.), Muthaiah Venkatachalam (Intel Corporation) ·

    Cascadia:Eleven AI PC 上的 Resident 975B MoE 推理

    arXiv:2610.07219v1 Announce Type: new Abstract: Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia's reside…

  2. arXiv cs.AI TIER_1 English(EN) · Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager ·

    DySCo:用于深度同步批处理的协作边缘-云 LLM 推理的动态分片

    arXiv:2610.08268v1 Announce Type: cross Abstract: Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their hi…

  3. arXiv cs.AI TIER_1 English(EN) · Zhi-Kai Chen, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye ·

    SchemaFill:通过槽位并行推测解码实现高效 LLM 工具调用

    arXiv:2610.07086v1 Announce Type: cross Abstract: LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, poten…

  4. arXiv cs.AI TIER_1 English(EN) · Chence Yang, Ningxi Cheng, Arash Akbari, Qitao Tan, Qingchan Zhu, Ci Zhang, Changdi Yang, Yanzhi Wang, Wei Niu, Jinhui Wang, Jin Lu, Geng Yuan ·

    BitNest:位嵌套的推测解码,用于内存高效的大模型推理加速

    arXiv:2610.02800v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation…

  5. arXiv cs.AI TIER_1 English(EN) · Indranil Halder, Cengiz Pehlevan ·

    揭秘LLM-as-a-Judge:用于推理时扩展的分析可处理模型

    arXiv:2512.19905v3 Announce Type: replace-cross Abstract: Recent developments in large language models have shown advantages in reallocating a notable share of computational resource from training time to inference time. However, the principles behind inference time scaling are n…

  6. arXiv cs.LG TIER_1 English(EN) · Zhihui Gao, Tingjun Chen, Dirk Englund ·

    AIR-LLM:通过射频计算在无线电上广播 AI 权重,实现无内存边缘 LLM 推理

    arXiv:2610.00465v1 Announce Type: cross Abstract: Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spen…

  7. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jiangsu Du ·

    EdgeAgent:在 CPU-GPU 统一内存架构上为终端用户多智能体系统编排设备端 LLM 推理

    Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU…

  8. arXiv cs.AI TIER_1 English(EN) · Devleena Das, Rajeev Patwari, Vikram Kumar Bukka, Nithin Kumar Guggilla, Elliott Delaye, Ashish Sirasao ·

    MINCE:通过少模型蒙特卡洛校准缩小 LLM 评估数据集

    arXiv:2606.22826v2 Announce Type: replace Abstract: Evaluating LLMs across many model variants---quantized, fine-tuned, or deployment-specific---requires running large benchmarks repeatedly, a process that can take tens of hours per model on edge hardware such as NPUs. Existing s…

  9. arXiv cs.LG TIER_1 English(EN) · Jitai Hao, Quansheng Gu, Qiang Huang, Jun Yu ·

    SparseEngine: Sparse-First 推理引擎

    arXiv:2609.39068v1 Announce Type: new Abstract: Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with…

  10. arXiv cs.AI TIER_1 English(EN) · Mukund Agarwalla, Chih-Jen Lin ·

    TopK-Guided:高效LLM推理的自适应、预算感知激活稀疏性

    arXiv:2610.01763v1 Announce Type: new Abstract: Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs:…

  11. arXiv cs.AI TIER_1 English(EN) · Junxuan Li, Arko Mukherjee, Soumyabrata Pal ·

    LLM 裁判在稀疏重叠下的验证:从推理到设计

    arXiv:2609.31857v2 Announce Type: replace Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deplo…

  12. arXiv cs.AI TIER_1 English(EN) · Ian Arawjo ·

    如何对LLM裁判进行统计分析并信任结果:使用evalstats进行小样本AI评估的校准推断

    arXiv:2609.35815v1 Announce Type: cross Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such clai…

  13. arXiv cs.LG TIER_1 English(EN) · Yufan Zhang, Sagnik Mukherjee, Hao Peng ·

    从 Fisher 视角理解 LLM 参数更新稀疏性

    arXiv:2609.36262v1 Announce Type: new Abstract: Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and s…

  14. arXiv cs.AI TIER_1 English(EN) · Jiyoung Park, Hankyu Jang, Changseok Song, Wookeun Jung ·

    TIDE: 用于自改进LLM推理的时间增量草稿引擎

    arXiv:2602.05145v2 Announce Type: replace-cross Abstract: Speculative decoding can substantially accelerate LLM inference, but realizing its benefits in practice is challenging due to evolving workloads. We present TIDE (Temporal Incremental Draft Engine), a serving-engine-native…

  15. Together AI blog TIER_1 English(EN) ·

    通过 IBM Cloud 和 NVIDIA 扩展我们的企业推理能力

    Enterprises can now run open models at production scale on a dedicated B300 inference cluster, built by Together AI, IBM Cloud, and NVIDIA

  16. Medium — MLOps tag TIER_1 English(EN) · Muharrem Bozkuş ·

    停止购买 GPU:通过更智能的路由和 KV 缓存管理来扩展 LLM 推理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/stop-buying-gpus-scale-llm-inference-with-smarter-routing-and-kv-cache-management-785c6be30911?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1376/0*jcsdOfpvoxM0W…

  17. Medium — MLOps tag TIER_1 English(EN) · Harshit Dawar ·

    大型语言模型(LLM)推理如何工作?详细解释,包括所有阶段!

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://harshitdawar.medium.com/how-does-llm-inference-work-explained-in-detail-including-all-its-phases-bafc98e8fe5f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/1*O92G2f-pdsPMmb46…

  18. Medium — MLOps tag TIER_1 English(EN) · Zeenatriaz ·

    MLOps终极指南:使用vLLM和AMD ROCm优化LLM推理

    <div class="medium-feed-item"><p class="medium-feed-snippet">A production-ready blueprint for high-throughput machine learning systems engineering, advanced hardware acceleration, and zero-trust&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@zeenatriaz468/ult…

  19. dev.to — LLM tag TIER_1 English(EN) · Alex Chen ·

    我比较了 2026 年的 5 个 GPU 云用于 LLM 推理——我发现了什么

    <h1> I Compared 5 GPU Clouds for LLM Inference in 2026 — Here's What I Found </h1> <p><em>Researched October 2026. Prices from public pricing pages and third-party trackers.</em></p> <p>Running LLMs in production gets expensive fast. I spent a week comparing GPU cloud providers t…

  20. r/LocalLLaMA TIER_1 English(EN) · /u/Distinct-Pie2389 ·

    LLM推理仪表板

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wxm2og/llm_inference_dashboard/"> <img alt="LLM Inference Dashboard" src="https://preview.redd.it/i0wcup9bphth1.jpg?width=140&amp;height=80&amp;auto=webp&amp;s=a268356e31d6b4f5e4aa5f55dbce03a1c5798caa" title=…

  21. r/LocalLLaMA TIER_1 English(EN) · /u/carteakey ·

    过拟合推理引擎的兴起

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/"> <img alt="The Rise of Overfit Inference Engines" src="https://external-preview.redd.it/aNcJUYeEP8Q3f5LSIpNIZjjRXbxAsnMQfvsJVWEZJys.png?width=640&amp;crop=smart&…

  22. dev.to — LLM tag TIER_1 English(EN) · Saptarshi Paul ·

    使用 WebGPU 在浏览器中进行 LLM 推理:2026 年实地指南

    <p>Nearly every AI feature in a web app today is a round trip: the page collects text, sends it to a hosted model, and streams tokens back over the network. That design puts a per-token bill, a network hop, and someone else's data retention policy between your user and a text box…

  23. r/LocalLLaMA TIER_1 English(EN) · /u/Postmodern_Plunger ·

    面向小白的推理工程

    <!-- SC_OFF --><div class="md"><p>Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing ru…

  24. dev.to — LLM tag TIER_1 English(EN) · Yuri Pocepaev ·

    一个模型,多种角色:无需训练更多模型即可实现 LLM 推理的专业化

    <p>I ran the same ten tool-selection tasks with four different tool catalogues against one shared model backend: <strong>40 requests in total</strong>.</p> <p>With five tools, the requests used <strong>6,208 input tokens</strong> in total. With fifty tools, they used <strong>38,2…

  25. Mastodon — mastodon.social TIER_1 English(EN) · adityahalderdev ·

    面向开发者的 Qwen3.8-Flash-Next 本地推理指南:硬件选型、本地验证、OpenAI 兼容客户端及性能基准测试

    A developer-first guide to Qwen3.8-Flash-Next local inference with Strata: hardware sizing, localhost verification, an OpenAI-compatible client, and benchmark caveats. Vendor and community numbers are labeled. https:// codereportglobal.indevs.in/art icles/qwen38-flash-next-local-…