PulseAugur
EN
LIVE 12:00:47

HauhauCS releases faster, uncensored Gemma 4 models with MTP

HauhauCS has released new versions of their Gemma 4 models, including 26B-A4B and 31B variants, which are uncensored and feature multi-token prediction (MTP) for increased speed. The 26B-A4B model is an MoE architecture offering approximately 35% faster generation, while the 31B model is a dense architecture providing about 53% faster generation. Additionally, a 12B Gemma 4 QAT Uncensored Balanced model with MTP has been released, boasting around a 60% speed boost. These models are designed for creative writing and role-playing, with the 26B-A4B recommended for general use and the 31B for users with sufficient VRAM. AI

IMPACT These releases offer improved performance and accessibility for local LLM users, potentially increasing adoption of these models for creative and general tasks.

RANK_REASON Release of quantized, fine-tuned models from a community developer, not a frontier lab.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

HauhauCS releases faster, uncensored Gemma 4 models with MTP

COVERAGE [2]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/hauhau901 ·

    Gemma4-26B-A4B & 31B-QAT Uncensored Balanced are out with MTP (35% & 53% speed boost)!

    <!-- SC_OFF --><div class="md"><p>First of all, I'm stoked to announce <strong>we are almost at 20 million downloads on HF!</strong> (counted only on my own account, no duplicates/quants/finetunes/etc) <strong>and almost 5000 members on Discord!</strong></p> <p>Two releases this …

  2. r/LocalLLaMA TIER_1 English(EN) · /u/hauhau901 ·

    Gemma4-12B-QAT Uncensored Balanced is out with MTP (~60% speed boost)!

    <!-- SC_OFF --><div class="md"><p>First of all, I'm stoked to announce <strong>we are almost at 20 million downloads on HF!</strong> (counted only on my own account, no duplicates/quants/finetunes/etc) <strong>and almost 5000 members on Discord!</strong></p> <p><a href="https://h…