PulseAugur
中
实时 16:01:06
Español(ES) Cómo correr un modelo de 180B sin GPU: POCKET-Darwin-180B en GGUF con llama.cpp

180B 参数 AI 模型通过量化在消费级硬件上运行

名为 POCKET-Darwin-180B-GGUF 的 Darwin-180B-RSI 模型的量化版本已发布,使其无需专用 GPU 即可在消费级硬件上运行。这个 111 GB 的模型采用了专家混合(MoE)架构,这意味着其 1800 亿参数中只有一小部分在每个 token 生成时处于激活状态。该模型在具有足够 RAM 的 CPU 上可以达到每秒 21 个 token 的速度,或者在 RAM 较少的笔记本电脑上通过利用 SSD 加载权重来获得较慢的速度,同时保持原始模型的精度。 AI

影响 使得在消费级硬件上运行大型语言模型成为可能,减少了对云 GPU 的依赖。

排序理由 发布了用于本地推理的量化、开源模型。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

180B 参数 AI 模型通过量化在消费级硬件上运行

本文如何被排名

Signal score
22 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
发布了用于本地推理的量化、开源模型。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 Español(ES) · 김민식/학생 ·

    如何在没有 GPU 的情况下运行 180B 模型:使用 llama.cpp 在 GGUF 中运行 POCKET-Darwin-180B

    <h2> TL;DR </h2> <p>POCKET-Darwin-180B-GGUF es la versión cuantizada del modelo abierto Darwin-180B-RSI, pensada para correr sin GPU. Puntos clave para quien va a desplegarlo:</p> <ul> <li> <strong>Pesa 111 GB</strong> en GGUF (el original en BF16 ocupa 360 GB) y se ejecuta con <…