PulseAugur
EN
LIVE 04:51:27

Ling-3.0-flash model shows narrow speed range across quantizations on DGX Spark

A user on Reddit shared benchmark results for the Ling-3.0-flash model, highlighting its performance on a DGX Spark system. The results show a narrow speed range of 32 to 40 tokens per second across different quantization levels, with the Q5_K_M quantization being both the fastest and near-lossless. This performance is significantly better than DeepSeek V4 Flash on the same hardware, with Ling-3.0-flash achieving approximately 2.4 times the speed. AI

IMPACT Demonstrates efficient performance scaling across quantization levels for large language models on specialized hardware.

RANK_REASON User-generated benchmark results for a specific model and hardware configuration. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Ling-3.0-flash model shows narrow speed range across quantizations on DGX Spark

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/AcanthisittaOk1699 ·

    Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vlmun8/ling30flash_quant_ladder_on_one_dgx_spark_the/"> <img alt="Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band" src="https://preview.redd.it/lko1z9fxyrih1.jpg?wi…