Google DeepMind has introduced agentic video understanding capabilities to its Gemini models, allowing them to analyze videos more efficiently. Instead of processing entire files, Gemini now reasons across transcripts, audio, and frames, dynamically adjusting frame rates to pinpoint crucial moments. This enhancement is particularly beneficial for long-form content and results in up to 88% fewer tokens used, with efficiency gains most significant for extended recordings. The feature is rolling out to Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via API in Google AI Studio and will soon be available in the Gemini app. AI
IMPACT Enhances efficiency for video analysis in AI models, potentially reducing costs and improving performance for long-form content.
RANK_REASON This is a feature update to existing models, not a new model release or significant research breakthrough.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →