PulseAugur
中
实时 10:39:04
English(EN) Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

Inkling-Small 276B-A12B 模型针对低内存消费级硬件进行了优化

Inkling-Small 276B-A12B 模型(约有 120 亿活跃参数)的新转换版本已针对在内存小于 10GB 的消费级硬件上运行进行了优化。基准测试显示,该模型在生成时速度约为每秒 2.9 个 token,但长提示预填充时间被指出是一个问题。这一发展是使大型混合专家(MoE)模型在标准硬件上可访问的持续努力的一部分。 AI

影响 使得在消费级硬件上运行大型 MoE 模型成为可能,从而可能拓宽其可访问性和用例。

排序理由 该项目描述了一个特定模型转换的优化和发布,适用于消费级硬件,属于工具范畴。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Inkling-Small 276B-A12B 模型针对低内存消费级硬件进行了优化

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一个特定模型转换的优化和发布,适用于消费级硬件,属于工具范畴。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
63 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Blahblahblakha ·

    Inkling-Small 276B-A12B 在 <10gb 内存上以 ~2.9 tok/s 运行

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vgfuyg/inklingsmall_276ba12b_at_29_toks_on_10gb_memory/"> <img alt="Inkling-Small 276B-A12B at ~2.9 tok/s on &lt;10gb memory" src="https://external-preview.redd.it/dTMzaGk2dG5vbGhoMctJqbS-x_J0wEm4fbWvJ83dxP-C…