A community project has demonstrated that a consumer-grade RTX 4090 GPU can achieve 100 trillion tokens per second when running the Qwen 3.8 Flash Next large language model. This feat was accomplished through aggressive int4 quantization, a speculative decoding pipeline using a smaller drafter model, and a fused inference stack optimized with TensorRT-LLM. The achievement significantly lowers the cost of running large LLMs, making them more accessible to researchers and power users, and challenges Nvidia's marketing of its high-end H100 GPUs for data-center-only performance. AI
IMPACT Significantly lowers the cost of LLM inference, democratizing access to large models for researchers and power users.
RANK_REASON Demonstration of consumer hardware achieving data-center-level LLM inference performance. [lever_c_demoted from significant: ic=1 ai=1.0]
- AMD Ryzen 9 7950X3D
- CUDA
- GitHub
- Nvidia
- NVIDIA H100
- Quantize the Drafter
- Quantize the Target
- Qwen 3.8 Flash Next
- RTX 4090
- Strata
- TensorRT-LLM
- Ubuntu 24.04
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →