PulseAugur
EN
LIVE 22:18:32

Custom CUDA kernel boosts Qwen3.8-27B performance by 1.9x on RTX 3090

A user has developed a custom CUDA megakernel for the Qwen3.8-27B model, significantly boosting its performance on an RTX 3090 graphics card. This new kernel achieves speeds 1.4x to 1.9x faster than standard llama.cpp implementations for tasks like code generation and prompt processing. The optimization involves running speculative decoding cycles within a single kernel launch, reducing overhead. While currently tailored for a specific model quantization and hardware, the kernel is designed as an OpenAI-compatible server replacement. AI

IMPACT Demonstrates potential for significant inference speedups on consumer hardware through custom kernel development.

RANK_REASON User-developed optimization for an open-source model. [lever_c_demoted from research: ic=1 ai=0.7]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Custom CUDA kernel boosts Qwen3.8-27B performance by 1.9x on RTX 3090

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-developed optimization for an open-source model. [lever_c_demoted from research: ic=1 ai=0.7]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Adorable_Weakness_39 ·

    Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel

    <!-- SC_OFF --><div class="md"><p>I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length. The results and code are below:</p> <p>Results (same hardware, my…