Google DeepMind has released DiffusionGemma, an open-source AI model that generates text in parallel blocks rather than sequentially, a departure from traditional token-by-token generation. This block-diffusion approach allows the model to refine entire segments of text simultaneously, leading to significantly faster inference speeds, reportedly up to four times faster than comparable Gemma models on a single NVIDIA H100. The DiffusionGemma model is built on the Gemma 4 26B architecture, features a 256K context window, supports over 140 languages, and can process text, image, and video inputs, all under the permissive Apache 2.0 license. AI
IMPACT Introduces a new parallel generation paradigm that could significantly speed up LLM inference, shifting focus from hardware to algorithmic innovation.
RANK_REASON Frontier-lab model release with novel generation technique. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Apache Software License 2.0
- DiffusionGemma
- Gemma
- Gemma 4: 26b
- Google DeepMind
- Groq
- NVIDIA
- NVIDIA H100
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →