PulseAugur
中
实时 20:52:55
English(EN) Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)

Basalt 推理引擎为 Qwen3.8 Flash-Next 提供 2.6 倍速度提升

一个名为 Basalt 的新推理引擎已被开发出来,为特定的语言模型提供了显著的速度提升。该引擎是 Strata 的一个分支,针对 Qwen3.8 Flash-Next 和 Blackwell 架构进行了优化,吞吐量达到了其前身的 2.6 倍。Basalt 支持双 GPU,并配备了一个自定义视觉编码器,该编码器比现有实现快得多,同时还支持多用户实时并发。 AI

影响 为特定的 LLM 配置提供了显著的性能提升,有可能提高兼容硬件用户的本地推理速度。

排序理由 这是现有推理引擎(Strata)的一个分支,针对特定硬件和模型进行了性能优化,而不是新颖的前沿模型发布或重大的行业范围发展。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Basalt 推理引擎为 Qwen3.8 Flash-Next 提供 2.6 倍速度提升

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是现有推理引擎(Strata)的一个分支,针对特定硬件和模型进行了性能优化,而不是新颖的前沿模型发布或重大的行业范围发展。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/jesdga95 ·

    Basalt: Flash-Next 在 5090 + 5060 Ti 上实现 665 tok/s 结构化、354 prose(2.6x Strata)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x223ai/basalt_flashnext_at_665_toks_structured_354_prose/"> <img alt="Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)" src="https://preview.redd.it/txboeb3odjuh1.png?wi…