A user on r/LocalLLaMA shared performance benchmarks for several new large language models, including DeepSeek V4 Flash, Qwen3.8 Flash Next, Qwen3.8-27B, and Qwen3.6-35B-A3B. The tests were conducted on NVIDIA DGX Spark hardware and focused on metrics like runtime, served context length, and delivered tokens per second. The Qwen3.8 Flash Next model, particularly with a 'medium' setting, achieved the highest score of 22/24, while its 'low' setting offered a better everyday balance. DeepSeek V4 Flash demonstrated a large output token generation capability but at a significantly longer runtime. AI
IMPACT Provides practical performance data for users running local LLMs on DGX hardware, highlighting trade-offs between speed, context length, and output quality.
RANK_REASON User-conducted benchmarks of multiple LLMs on specific hardware. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →