PulseAugur
实时 21:03:24
English(EN) Glimmer: 233.4 tps on 5090 with Dflash

Glimmer LLM 在 5090 GPU 上实现 233.4 tps,达到 256k 上下文

一款名为 Glimmer 的新模型展示了令人印象深刻的性能,在配备 Dflash 的 5090 GPU 上实现了 233.4 tps。用户报告称,Glimmer 可以在 24GB VRAM 上轻松达到 256k 的上下文窗口,这是 Qwen 等模型难以实现的壮举。这一发展在本地 LLM 社区中引起了极大的兴奋。 AI

影响 为具有大型上下文窗口的更强大、更高效的本地 LLM 部署提供了潜力。

排序理由 新模型发布,具有特定的性能指标和功能。[lever_c_demoted from frontier_release: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Glimmer LLM 在 5090 GPU 上实现 233.4 tps,达到 256k 上下文

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/YetAnotherAnonymoose ·

    Glimmer: 233.4 tps on 5090 with Dflash

    <!-- SC_OFF --><div class="md"><p>That's insane yo! I haven't had a chance yet to test it on my 4090 at home but it sounds so promising. And read here that 256k CTX is easily reachable on 24gb unlike Qwen. Super excited!</p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="…