PulseAugur
中
实时 20:53:20
English(EN) Part 4 of 4, How AI Actually Works: what does a language model really cost to run? I measured Qwen3-8B at 16, 8, 4 and 2 bits on an M4 Pro MacBook with llama.cp

Qwen3-8B 语言模型成本分析显示 4 比特精度几乎免费

一项技术分析探讨了在 M4 Pro MacBook 上使用 llama.cpp 运行 Qwen3-8B 语言模型的运营成本。该研究在各种比特精度(16、8、4 和 2 比特)下测量了性能,评估了文件大小、困惑度、读写速度以及高达 65,000 个 token 的 KV 缓存增长。结果表明,四比特精度运行成本几乎为零,而二比特精度在尺寸和速度方面有优势,但以牺牲性能为代价。 AI

影响 提供了关于在不同精度级别下运行大型语言模型的实际成本和权衡的见解。

排序理由 模型性能和成本的技术分析。[lever_c_从研究中降级:ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3-8B 语言模型成本分析显示 4 比特精度几乎免费

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
模型性能和成本的技术分析。[lever_c_从研究中降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 English(EN) · R4TSQ ·

    第四部分(共四部分),人工智能如何运作:运行语言模型实际成本是多少?我在 M4 Pro MacBook 上以 16、8、4 和 2 位精度测量了 Qwen3-8B 在 llama.cp 上的表现

    Part 4 of 4, How AI Actually Works: what does a language model really cost to run? I measured Qwen3-8B at 16, 8, 4 and 2 bits on an M4 Pro MacBook with llama.cpp: file size, perplexity, reading and writing speed, and the KV cache growing to 65,000 tokens. Four bits was nearly fre…