研究人员正在开发新颖的方法来提高大型语言模型(LLM)推理的效率。一种方法AIR-LLM提出通过射频广播LLM权重,使边缘设备无需存储权重即可进行推理,从而显著降低能耗和通信时长。其他研究则侧重于通过稀疏性技术优化推理,例如SparseEngine和TopK-Guided,它们通过选择性地处理模型的一部分来降低内存和计算成本。此外,正在开发MINCE和evalstats等新框架,通过缩小数据集和提高LLM评判分数统计分析的可靠性来简化LLM评估。
AI
arXiv:2610.07219v1 Announce Type: new Abstract: Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia's reside…
arXiv cs.AI
TIER_1English(EN)·Jingpo Xu, Paul Joe Maliakel, Ivona Brandic, Shashikant Ilager·
arXiv:2610.08268v1 Announce Type: cross Abstract: Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their hi…
arXiv:2610.07086v1 Announce Type: cross Abstract: LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, poten…
arXiv:2610.02800v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation…
arXiv:2512.19905v3 Announce Type: replace-cross Abstract: Recent developments in large language models have shown advantages in reallocating a notable share of computational resource from training time to inference time. However, the principles behind inference time scaling are n…
arXiv:2610.00465v1 Announce Type: cross Abstract: Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spen…
Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU…
arXiv:2606.22826v2 Announce Type: replace Abstract: Evaluating LLMs across many model variants---quantized, fine-tuned, or deployment-specific---requires running large benchmarks repeatedly, a process that can take tens of hours per model on edge hardware such as NPUs. Existing s…
arXiv cs.LG
TIER_1English(EN)·Jitai Hao, Quansheng Gu, Qiang Huang, Jun Yu·
arXiv:2610.01763v1 Announce Type: new Abstract: Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs:…
arXiv:2609.31857v2 Announce Type: replace Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deplo…
arXiv:2609.35815v1 Announce Type: cross Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such clai…
arXiv:2609.36262v1 Announce Type: new Abstract: Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and s…
arXiv cs.AI
TIER_1English(EN)·Jiyoung Park, Hankyu Jang, Changseok Song, Wookeun Jung·
arXiv:2602.05145v2 Announce Type: replace-cross Abstract: Speculative decoding can substantially accelerate LLM inference, but realizing its benefits in practice is challenging due to evolving workloads. We present TIDE (Temporal Incremental Draft Engine), a serving-engine-native…
<h1> I Compared 5 GPU Clouds for LLM Inference in 2026 — Here's What I Found </h1> <p><em>Researched October 2026. Prices from public pricing pages and third-party trackers.</em></p> <p>Running LLMs in production gets expensive fast. I spent a week comparing GPU cloud providers t…
<p>Nearly every AI feature in a web app today is a round trip: the page collects text, sends it to a hosted model, and streams tokens back over the network. That design puts a per-token bill, a network hop, and someone else's data retention policy between your user and a text box…
<!-- SC_OFF --><div class="md"><p>Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing ru…
<p>I ran the same ten tool-selection tasks with four different tool catalogues against one shared model backend: <strong>40 requests in total</strong>.</p> <p>With five tools, the requests used <strong>6,208 input tokens</strong> in total. With fifty tools, they used <strong>38,2…
A developer-first guide to Qwen3.8-Flash-Next local inference with Strata: hardware sizing, localhost verification, an OpenAI-compatible client, and benchmark caveats. Vendor and community numbers are labeled. https:// codereportglobal.indevs.in/art icles/qwen38-flash-next-local-…