Google has introduced Gemma 4, a new family of open-weight, multimodal language models. These models feature dense and Mixture-of-Experts architectures, with parameter counts ranging from 2.3B to 31B. Gemma 4 includes enhanced vision and audio encoders, and a novel encoder-free architecture for its 12B model that processes raw audio and image patches. The models also incorporate a "thinking mode" for generating reasoning traces before providing responses, alongside improvements in inference speed, memory efficiency, and long-context capabilities. AI
IMPACT Sets new SOTA on STEM and multimodal benchmarks, potentially challenging existing open models.
RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gemma
- Gemma 4
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →