PulseAugur
实时 00:23:38
English(EN) NInfer fork: 555k context@fp4 for 5090 with YARN, reliable kv host cacheing, monitoring, jinja, opened model support

NInfer 分叉将 LLM 上下文提升至 555k,采用 4 位 KV 缓存

NInfer 项目的一个分叉已被开发出来,为大型语言模型的上下文长度和内存管理带来了显著的改进。该分叉采用自定义的 4 位 KV 缓存,可在不损失质量的情况下将 VRAM 使用量减少 45%,并通过 LongBench 和 AIME 等基准测试得到验证。它还将 Qwen 等模型的上下文窗口扩展到 555k token 以上,并预计在高 端 GPU 上可达 800 万 token。该项目增强了多级前缀重用,并实现了一个强大的主机 KV 缓存安全网,以确保并发会话的稳定性能。 AI

影响 提高了本地 LLM 部署的效率和上下文处理能力,有可能在消费级硬件上实现更复杂的任务。

排序理由 这是现有工具的一个分叉,具有新功能,而不是前沿模型发布或重大的行业事件。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

NInfer 分叉将 LLM 上下文提升至 555k,采用 4 位 KV 缓存

本文如何被排名

Signal score
8 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是现有工具的一个分叉,具有新功能,而不是前沿模型发布或重大的行业事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Lumpy-Comedian-1027 ·

    NInfer 分叉:5090 上的 555k 上下文@fp4,支持 YARN、可靠的 kv 主机缓存、监控、jinja、开放模型

    <!-- SC_OFF --><div class="md"><p>Hiya,</p> <p>NInfer is amazng for Qwen, but lacking for real-world-use. As adoption of issues/pr's was not really what I needed, I created a fork and hit it for this week with 3 concurrent claude code session until it didn't break any longer. Hop…