Researchers have introduced DiffusionGemma, an experimental open-weight language model designed for high-speed text generation. Unlike traditional autoregressive models that process tokens sequentially, DiffusionGemma refines blocks of 256 tokens in parallel using discrete diffusion. This model is derived from the Gemma 4 mixture-of-experts model and achieves approximately 1,500 output tokens per second on a single NVIDIA H100 GPU, significantly outperforming conventional methods. AI
IMPACT Establishes a new Pareto frontier for generation speed and model capability, potentially accelerating AI applications requiring rapid text output.
RANK_REASON Publication of a technical report detailing a new experimental language model on arXiv.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →