A new port of the NInfer inference engine for Windows has been released, optimized for the RTX 4090 graphics card. This port allows users to run the Qwen 3.8 27B model with performance ranging from 60-100 tokens per second. The engine supports context lengths of up to 100-150K tokens, depending on the specific configuration. AI
IMPACT Enables local execution of large language models on consumer hardware, improving accessibility for developers and researchers.
RANK_REASON This is a software port and optimization for existing hardware and models, not a novel model release or research breakthrough.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →