A software engineer has developed Ninfer 4080, a system designed to run the ISTA-DASLab-Qwen-3.8-27B-GSQ model on an RTX 4080 GPU with 16GB of memory. This new system aims to significantly improve prefill and token generation speeds, achieving up to 2720 tokens/sec for prefill and 262 tokens/sec for generation at a 100k context length. The project, shared on GitHub, utilizes DFlash2 speculative decoding and aims to optimize hardware utilization beyond general-purpose inference engines. AI
IMPACT Enables running larger context models on consumer-grade hardware, potentially lowering the barrier for advanced LLM experimentation.
RANK_REASON This is a user-developed tool for running existing models on specific hardware, not a release from a frontier lab.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →