PulseAugur
中
实时 05:17:30
English(EN) [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

Gemma 2B 在 5KB x86-64 汇编引擎上运行

一位开发者为 Gemma 2B 语言模型创建了一个高度优化的推理引擎,完全用 x86-64 汇编编写。该引擎拥有仅 5.2 KB 的最小二进制占用空间,并在四核 i5 CPU 上使用 FP16 精度实现了约 4.6 个 token/秒。该项目名为 PULSAR-ASM,旨在探索运行现代 Transformer 的裸机要求,并作为资源受限微控制器的参考,摒弃了 C/C++ 或 PyTorch 等传统运行时。 AI

影响 展示了 LLM 在资源受限硬件上的极端优化潜力。

排序理由 开发者为现有模型创建的最小占用空间优化。 [lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Gemma 2B 在 5KB x86-64 汇编引擎上运行

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
开发者为现有模型创建的最小占用空间优化。 [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/tom_tsai28 ·

    [讨论] Gemma-2B 的 5KB 純 x86-64 汇编引擎 (FP16, CPU 速度 4.6 toks/s)

    <!-- SC_OFF --><div class="md"><p>Hi everyone,</p> <p>Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM.</p> <p>Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entire…