PulseAugur
EN
LIVE 08:30:56

Qwen3.6-27B model optimized for V100 GPUs hits 366 t/s

A developer has optimized the Qwen3.6-27B model for NVIDIA V100 GPUs, achieving up to 366 tokens per second in specific benchmarks. This optimization, named "v100-skinny," focuses on creating fast paths for NVFP4 weights and enabling efficient deep speculation on SM70 architecture. While the peak performance is noted for specific extraction tasks, practical generation speeds are around 240 tokens per second for structured data like JSON and 200 tokens per second for code generation. AI

IMPACT Demonstrates significant performance gains for specific hardware, potentially enabling faster local inference for certain models.

RANK_REASON Developer-led optimization of an existing model for specific hardware, not a frontier lab release. [lever_c_demoted from research: ic=1 ai=0.7]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.6-27B model optimized for V100 GPUs hits 366 t/s

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Developer-led optimization of an existing model for specific hardware, not a frontier lab release. [lever_c_demoted from research: ic=1 ai=0.7]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
57 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Simple_Library_2700 ·

    366 t/s Qwen3.6 27B NVFP4 on v100s

    <!-- SC_OFF --><div class="md"><p><strong>These are single stream numbers</strong></p> <p>Following on from my previous post about v100s (<a href="https://www.reddit.com/r/LocalLLaMA/comments/1tmyln6/1000_tps_generation_on_qwen36_27b_with_v100s/">here</a>) and inspired by this co…