Ollama has released version v0.33.3-rc2, introducing support for image and audio input for Gemma4 models. This update leverages the MLX engine to process multimodal inputs, with images being handled by both transformer tower and unified embedder architectures, and audio processed through a dedicated intake that accepts WAV bytes and audio transcriptions. The system intelligently routes modalities to compatible checkpoints, ensuring that models without specific audio or vision capabilities function as text-only and reject unsupported input types. AI
IMPACT Enhances multimodal capabilities for local AI model deployment.
RANK_REASON This is a software release for a tool that integrates AI models, not a release of a frontier model itself.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →