A user shared benchmark results for the DeepSeek-V4-Flash-0731-UD-Q3_K_XL model, tested on a setup with three NVIDIA RTX 3090 GPUs. The results indicate a tokens per second rate of 116.04 for the pp512 test and 7.71 for the tg128 test. The user noted that further optimization might be possible but has not yet achieved better performance. AI
IMPACT Provides performance data for a specific model configuration, useful for users with similar hardware.
RANK_REASON User-generated benchmark results for a specific model quantization. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →