PulseAugur
实时 03:02:35
English(EN) NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM

NInfer 更新使 RTX 4090 上的 Qwen 3.8-27B 支持 350K token 上下文

一位用户更新了他们 fork 的 NInfer(一个用于在本地运行大型语言模型的工具),以支持 Qwen 3.8-27B 模型。此次更新允许在单个 RTX 4090 GPU 上实现高达 250-350K token 的上下文窗口,而无需系统 RAM。这些优化还使得某些工作负载的生成速度达到每秒 80-160 token。 AI

影响 为本地 LLM 部署实现更大的上下文窗口,可能提高复杂任务的性能。

排序理由 用户开发的软件更新,用于在消费级硬件上运行现有模型。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

NInfer 更新使 RTX 4090 上的 Qwen 3.8-27B 支持 350K token 上下文

报道来源 [2]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/UDPSendToFailed ·

    NInfer RTX 4090 for Qwen 3.8 27B 更新 - VRAM 中高达 250-350K token 上下文

    <!-- SC_OFF --><div class="md"><p>I've made some improvements to my fork of NInfer, adding rk2v4-e8 quant option for the KV cache, which can reach up to 250-350K tokens of context window depending on the configuration like vision, MTP, etc, on my single RTX 4090 without spilling …

  2. r/LocalLLaMA TIER_1 English(EN) · /u/UDPSendToFailed ·

    Ninfer for RTX 4090 and Qwen 3.8 27B

    <!-- SC_OFF --><div class="md"><p>I made a quick port for Windows based on ninfer-3090. It seems to be working for the most part, reaching about 60-100t/s and fits up to 100-150K tokens depending on the context with rk8v4.</p> <p>Tested with qwen3_8_27b.ninfer from <a href="https…