PulseAugur
中
实时 17:45:01
English(EN) Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

Nifer 推理引擎在 Qwen 3.6 35B 模型上实现 720t/s

一款名为 Nifer 的新推理引擎已发布,据报道在 Qwen 3.6 35B 模型上实现了 550-720 tokens/秒的速度。用户形容这种性能“简直是疯了”,它无需复杂的批处理或并行代理即可实现,并支持 250k tokens 的上下文窗口。该引擎专门针对 RTX 5090 GPU 进行了优化,可在 GitHub 上获取,最初支持 Linux,未来可能支持 Windows。 AI

影响 可能实现大型语言模型显著更快的本地推理。

排序理由 发布了一款针对特定硬件优化的新推理引擎。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Nifer 推理引擎在 Qwen 3.6 35B 模型上实现 720t/s

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
发布了一款针对特定硬件优化的新推理引擎。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
72 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/BringTea_666 ·

    Nifer 疯狂了。Qwen 3.6 35B 达到 700t/s(无思考)。专为 RTX5090 构建。支持完整的 250k 上下文。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v8a7wb/nifer_is_insane_700ts_with_qwen_36_35b_no/"> <img alt="Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too." src="https://external-preview.redd.it/…