A new inference engine called Basalt has been developed, offering significant speed improvements for specific large language models. This engine, a fork of Strata, is optimized for Qwen3.8 Flash-Next and Blackwell architecture, achieving up to 2.6 times the throughput of its predecessor. Basalt supports dual GPUs and features a custom vision encoder that is substantially faster than existing implementations, alongside real concurrency for multiple users. AI
IMPACT Offers a significant performance boost for specific LLM configurations, potentially improving local inference speeds for users with compatible hardware.
RANK_REASON This is a fork of an existing inference engine (Strata) with performance optimizations for specific hardware and models, rather than a novel frontier model release or significant industry-wide development.
- 2.6x Strata
- 354 prose
- 5090
- 665 tok/s
- basalt
- Blackwell
- flash-next
- llama.cpp
- MIT license
- Nvidia RTX 5060 Ti
- Qwen3.8
- Strata
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →