PulseAugur
中
实时 12:39:54
English(EN) llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp

llama.cpp 为 MoE 模型添加 GPU 缓存以提升性能

一个拉取请求已合并到 llama.cpp 项目,为混合专家 (MoE) 模型引入了 GPU 缓存。此增强功能允许不完全适合 VRAM 的专家保留在主机内存中,从而可能为 GPU 资源有限的用户带来显著的速度提升。 AI

影响 此优化可能使用户能够在消费级硬件上运行更大的 MoE 模型。

排序理由 这是对一个开源项目的代码贡献,它提高了特定类型模型的性能。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

llama.cpp 为 MoE 模型添加 GPU 缓存以提升性能

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是对一个开源项目的代码贡献,它提高了特定类型模型的性能。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/jacek2023 ·

    llama : 为保留在主机内存中的 MoE 专家添加 GPU 缓存 · am17an 的拉取请求 #29887 · ggml-org/llama.cpp

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x03xkc/llama_add_a_gpu_cache_for_moe_experts_kept_in/"> <img alt="llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp" src="https://external-previe…