PulseAugur
EN
LIVE 02:32:03

Nifer inference engine achieves 720t/s on Qwen 3.6 35B model

A new inference engine called Nifer has been released, reportedly achieving speeds of 550-720 tokens per second on a Qwen 3.6 35B model. This performance, described as "insane" by users, is achieved without complex batching or parallel agents, and supports a context window of 250k tokens. The engine is specifically optimized for the RTX 5090 GPU and is available on GitHub, though initially for Linux with potential for Windows builds. AI

IMPACT Potentially enables significantly faster local inference for large language models.

RANK_REASON Release of a new inference engine optimized for specific hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Nifer inference engine achieves 720t/s on Qwen 3.6 35B model

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/BringTea_666 ·

    Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v8a7wb/nifer_is_insane_700ts_with_qwen_36_35b_no/"> <img alt="Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too." src="https://external-preview.redd.it/…