PulseAugur
EN
LIVE 04:51:26

Qwen3.6-27B model optimized for V100 GPUs hits 366 t/s

A developer has optimized the Qwen3.6-27B model for NVIDIA V100 GPUs, achieving up to 366 tokens per second in specific benchmarks. This optimization, named "v100-skinny," focuses on creating fast paths for NVFP4 weights and enabling efficient deep speculation on SM70 architecture. While the peak performance is noted for specific extraction tasks, practical generation speeds are around 240 tokens per second for structured data like JSON and 200 tokens per second for code generation. AI

IMPACT Demonstrates significant performance gains for specific hardware, potentially enabling faster local inference for certain models.

RANK_REASON Developer-led optimization of an existing model for specific hardware, not a frontier lab release. [lever_c_demoted from research: ic=1 ai=0.7]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.6-27B model optimized for V100 GPUs hits 366 t/s

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Simple_Library_2700 ·

    366 t/s Qwen3.6 27B NVFP4 on v100s

    <!-- SC_OFF --><div class="md"><p><strong>These are single stream numbers</strong></p> <p>Following on from my previous post about v100s (<a href="https://www.reddit.com/r/LocalLLaMA/comments/1tmyln6/1000_tps_generation_on_qwen36_27b_with_v100s/">here</a>) and inspired by this co…