Summary

On September 1, 2026, Google introduced agentic video understanding for its latest Gemini models, letting the model dynamically search, scan, and inspect target segments across frames, audio, and transcripts instead of ingesting video at a fixed frame rate. Google reports up to 88% fewer tokens, up to 66% lower analysis cost, and up to 7% better accuracy, live via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

What changed

Google shipped agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, pairing model reasoning with native video tools to dynamically inspect segments rather than statically sampling frames.

Why it matters

Agentic, tool-driven video processing sharply lowers the cost and token footprint of long-form video analysis, which has been prohibitively expensive at fixed frame sampling. It makes video a more practical first-class input for agents and enterprise workflows and pressures rivals on multimodal cost-efficiency.

Evidence excerpt

Agentic video understanding pairs the model's core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts... up to 88% fewer tokens, up to 66% lower analysis cost, and up to 7% better accuracy.

Sources