Google adds agentic video understanding to Gemini Flash models, claiming up to 88% lower token use
- Google released agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, claiming up to 88% lower token consumption, 66% lower analysis costs, and 7% higher accuracy on video benchmarks.
- Instead of ingesting a video at a fixed frame rate, agentic video understanding lets Gemini dynamically search, scan, and inspect relevant visual frames, audio, and transcripts.
- Google says the feature supports sub-second moment retrieval, more accurate anomaly detection, and precise counting, with the largest efficiency gains expected for videos lasting 10 minutes to multiple hours.
- The capability is available now for uploaded videos and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.