A new model architecture, Qwen 3.8 Flash Next, has demonstrated the ability to run on a single RTX 4090 GPU at approximately 100 tokens per second. This is achieved by employing a sparse Mixture-of-Experts (MoE) design where only a fraction of the model's parameters are active per token. The system, named Strata, offloads less frequently used parameters to system RAM and utilizes NVMe SSDs for n-gram tables, significantly reducing the VRAM requirement. However, achieving these speeds necessitates a substantial amount of system RAM, effectively requiring workstation-level memory alongside the consumer GPU. AI
IMPACT This development highlights innovative techniques for running large models on consumer hardware, potentially lowering the barrier to entry for advanced AI applications.
RANK_REASON The item discusses a novel model architecture and its performance on consumer hardware, which is a research-oriented topic. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →