PulseAugur
EN
LIVE 18:59:15

DeepSeek-V4-Flash Model Benchmarked on 3x RTX 3090 GPUs

A user shared benchmark results for the DeepSeek-V4-Flash-0731-UD-Q3_K_XL model, tested on a setup with three NVIDIA RTX 3090 GPUs. The results indicate a tokens per second rate of 116.04 for the pp512 test and 7.71 for the tg128 test. The user noted that further optimization might be possible but has not yet achieved better performance. AI

IMPACT Provides performance data for a specific model configuration, useful for users with similar hardware.

RANK_REASON User-generated benchmark results for a specific model quantization. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

DeepSeek-V4-Flash Model Benchmarked on 3x RTX 3090 GPUs

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/consultkitapp ·

    DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3x3090 test results

    <!-- SC_OFF --><div class="md"><p>For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far.</p> <h1>Command</h1> <p>./llama-bench -m /home/user/llamacpp/modelsmain/unsloth/ds4/DeepSeek-V4-Flash-073…