Gemini Adds Agentic Video Understanding, Fetching Only the Segments It Needs
Google launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite on September 1. Instead of ingesting a video whole at a fixed frame rate, the model fetches only the moments and modalities it needs.
Google launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite on September 1, 2026 1.
- Current “static” processing ingests a video at a fixed frames-per-second rate (1 FPS by default, adjustable via the API). The agentic version dynamically searches, scans, and inspects target video segments across visual frames, audio, and transcripts 1
- Across standard video analysis benchmarks, Google reports analysis costs down by up to 66%, token consumption down by up to 88%, and accuracy up by up to 7% — all “up to” figures rather than averages 1
- It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, at standard Gemini API token pricing with no additional feature fee. Set processing to
"agentic"in the API configuration to turn it on 1
Google says the efficiency gains are most pronounced on long-form video, where static processing “forces developers to choose between high token costs or techniques that drop critical details” 1. It targets a familiar bottleneck — token bills that balloon the moment you hand a whole video to a model — and if you are already on Gemini 3.7 Flash, released in August, it is a one-line API config change to try.
Sources
- Introducing agentic video understanding with Gemini - Google official blog (September 1, 2026)
Was this article helpful?
Thank you!
Received. Thank you!