PulseAugur
EN
LIVE 19:53:14

Basalt inference engine offers 2.6x speed boost for Qwen3.8 Flash-Next

A new inference engine called Basalt has been developed, offering significant speed improvements for specific large language models. This engine, a fork of Strata, is optimized for Qwen3.8 Flash-Next and Blackwell architecture, achieving up to 2.6 times the throughput of its predecessor. Basalt supports dual GPUs and features a custom vision encoder that is substantially faster than existing implementations, alongside real concurrency for multiple users. AI

IMPACT Offers a significant performance boost for specific LLM configurations, potentially improving local inference speeds for users with compatible hardware.

RANK_REASON This is a fork of an existing inference engine (Strata) with performance optimizations for specific hardware and models, rather than a novel frontier model release or significant industry-wide development.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Basalt inference engine offers 2.6x speed boost for Qwen3.8 Flash-Next

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
This is a fork of an existing inference engine (Strata) with performance optimizations for specific hardware and models, rather than a novel frontier model release or significant industry-wide deve…
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/jesdga95 ·

    Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x223ai/basalt_flashnext_at_665_toks_structured_354_prose/"> <img alt="Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)" src="https://preview.redd.it/txboeb3odjuh1.png?wi…