Google has launched a new capability for its Gemini models that fundamentally changes how video analysis works: agentic video understanding. Instead of processing videos frame-by-frame at a fixed rate, the models can now actively decide which segments to watch, how fast to scan them, and whether to use frames, audio, or transcripts. The results are impressive – and most importantly: significantly cheaper.
Quick Facts
- Available for: Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via Gemini API, Google AI Studio, and Gemini Enterprise Agent Platform
- Savings: Token consumption reduced by up to 88%, costs cut by up to 66%, accuracy improved by up to 7%
- Most effective for: Long-form videos (10-minute tutorials to 90-minute lectures and multi-hour recordings)
- Launch: Available now for video uploads and YouTube videos
How the Old System Worked – and Why It Was Inefficient
Previously, Gemini processed videos in static mode: the model ingested the entire file at a fixed frame rate (default 1 FPS, adjustable). For long videos, this created a dilemma: either high token costs or quality loss through sampling. Especially with long-form content (lectures, tutorials, multi-hour recordings), critical details were lost or the bill became unaffordable.
The New Approach: Targeted Video Analysis
Agentic video understanding works differently. The model takes active control: it decides which video segments are relevant, scans them at variable speeds, leverages parallel data sources (frames, audio, transcripts), and jumps directly to interesting moments. This is similar to agentic vision for images – but for videos.
The results are impressive benchmarks:
| Metric | Improvement |
|---|---|
| Token consumption | up to 88% reduction |
| Costs | up to 66% savings |
| Accuracy | up to 7% improvement |
Particularly noteworthy: Gemini 3.7 Flash with agentic understanding now sits at the "accuracy-to-cost Pareto frontier" – meaning no other tested model offers better accuracy at the same cost or lower cost at the same accuracy.
New Use Cases Become Possible
The capability enables applications that were previously too expensive or too inaccurate: sub-second moment retrieval ("Find the exact moment the product is shown"), more precise anomaly detection ("When does something unusual happen?"), accurate counting ("How many people are in the video?"). For enterprises automating video analysis, this is a genuine capability upgrade.
Activation is simple via API – just set the configuration to "agentic" in Google AI Studio or the Gemini Enterprise Agent Platform. Works with video uploads and YouTube videos.
What This Means for Enterprises
For companies integrating video into their KI workflows – from media organizations to security to document analysis – the cost equation changes significantly. 66% cost savings with better quality makes video-KI applications suddenly more economical. Especially for longer content (training videos, surveillance footage, archive analysis), this could be a tipping point. However: the capability is currently limited to Google's models – organizations already invested in other platforms need to re-evaluate whether switching makes sense.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




