PulseAugur
EN
LIVE 21:04:19

NInfer optimized for RTX 4090, supports Qwen 3.8 27B model

A new port of the NInfer inference engine for Windows has been released, optimized for the RTX 4090 graphics card. This port allows users to run the Qwen 3.8 27B model with performance ranging from 60-100 tokens per second. The engine supports context lengths of up to 100-150K tokens, depending on the specific configuration. AI

IMPACT Enables local execution of large language models on consumer hardware, improving accessibility for developers and researchers.

RANK_REASON This is a software port and optimization for existing hardware and models, not a novel model release or research breakthrough.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

NInfer optimized for RTX 4090, supports Qwen 3.8 27B model

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/UDPSendToFailed ·

    Ninfer for RTX 4090 and Qwen 3.8 27B

    <!-- SC_OFF --><div class="md"><p>I made a quick port for Windows based on ninfer-3090. It seems to be working for the most part, reaching about 60-100t/s and fits up to 100-150K tokens depending on the context with rk8v4.</p> <p>Tested with qwen3_8_27b.ninfer from <a href="https…