Google adds agent-based video retrieval to Gemini Flash, claiming 88% lower token use
- Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite now use agent-based video analysis that selects relevant video segments rather than sampling every second at a fixed rate.
- Google reports up to 88% lower token use and 66% lower cost on 1H-VideoQA and LVBench, while accuracy rises slightly versus static video processing.
- The model can choose whether to retrieve frames, audio, or transcripts, then resample suspicious time windows at higher frame rates to find short cuts, state changes, anomalies, and repeated actions.
- The capability is available in the Gemini API for scene search in long videos; Google plans Gemini app and YouTube integration later.
- Google says its LongVideoBench results put Gemini 3.7 Flash highest in overall quality among tested systems, with the strongest accuracy-to-cost result.