PulseAugur
EN
LIVE 08:28:51

DeepSeek V4 Flash model runs at 100 tok/s on four consumer GPUs

A user has successfully configured and run the DeepSeek V4 Flash model, a large 144 GiB MoE model, on a setup of four NVIDIA RTX 3060 12GB GPUs. This configuration achieved approximately 100 tokens/s for prompt processing and 10 tokens/s for text generation, while maintaining a context window of up to 368,640 tokens. The user detailed specific llama.cpp parameters and hardware configurations, highlighting the impact of microbatch size and tensor placement on performance and VRAM usage. AI

IMPACT Demonstrates efficient deployment of large MoE models on consumer-grade hardware, potentially lowering barriers to entry for advanced AI experimentation.

RANK_REASON User-driven configuration and performance report for running a large model on consumer hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

DeepSeek V4 Flash model runs at 100 tok/s on four consumer GPUs

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/syscomua ·

    Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vrqf4f/running_deepseek_v4_flash_q4_k_xl_at_100_toks/"> <img alt="Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB" src="https://preview.redd.it/lav8jwie65kh1.jpeg?width=6…