PulseAugur
EN
LIVE 17:52:27

Muse Glimmer LLM fits on single RTX 3090 with 256k context

The Muse Glimmer large language model has been found to comfortably fit on a single RTX 3090 graphics card, supporting a 256k context window with full features like DFlash and mmproj. This performance is notable as other models like Qwen3.6-27B and Gemma-4-31B struggle to achieve similar context lengths on the same hardware. Muse Glimmer also demonstrates impressive speed, processing prompts at approximately 1400 tokens/sec and generating text between 64-124 tokens/sec, while accurately retrieving information from extreme context lengths. AI

IMPACT Enables running advanced LLMs with large context windows on more accessible hardware, potentially democratizing complex AI applications.

RANK_REASON User-driven benchmark and performance analysis of an LLM on consumer hardware. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Muse Glimmer LLM fits on single RTX 3090 with 256k context

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/coder543 ·

    Muse Glimmer ACTUALLY fits on a single RTX 3090

    <!-- SC_OFF --><div class="md"><p>I did some testing this morning, and I was surprised to find that Muse Glimmer actually comfortably fits on a single RTX 3090 with full context + DFlash + mmproj at Q4_K_XL, unlike Qwen3.6-27B and Gemma-4-31B.</p> <p>Muse Glimmer supports up to 2…