PulseAugur
中
实时 07:40:07

SlimWise 框架提升 MoE 模型服务效率

研究人员开发了 SlimWise,一个旨在提高专家混合(MoE)模型服务效率的新框架。SlimWise 将专家剪枝解耦,仅在解码阶段应用,而在预填充阶段使用完整模型。这种方法允许在解码过程中重用预填充生成的 KV 缓存,显著减少流量瓶颈。该框架还包括一个低成本的蒸馏阶段,以进一步弥合准确性差距。在 vLLM 中实现后,SlimWise 在 Qwen3.6-35B-A3B 模型上实现了高达 1.81 倍的解码吞吐量提升,同时仅进行了 50% 的专家剪枝,并保持了最小的准确性损失。 AI

影响 提高 MoE 模型服务效率,可能降低大型语言模型的推理成本和延迟。

排序理由 该集群包含一篇详细介绍用于优化 AI 模型服务的新的技术框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

SlimWise 框架提升 MoE 模型服务效率

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍用于优化 AI 模型服务的新的技术框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
10 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    SlimWise:解耦预填充和解码中的专家剪枝,实现高效MoE服务

    Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacri…