Introducing agentic video understanding with Gemini
We’re launching agentic video understanding across our latest Gemini models for improved accuracy and lower costs and token usage.
Google has introduced agentic video understanding for its Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models. This new feature allows the model to dynamically scan video segments, improving accuracy while reducing token usage by up to 88% and costs by up to 66%. Developers can start using this capability by configuring their API settings to "agentic" in Google AI Studio or the Gemini Enterprise Agent Platform.
Agentic video understanding enables Gemini to take an active, goal-directed role in determining what to watch, at what speed, and through which modality (frames, audio, or transcript), fetching only the moments and signals needed. This transformation in processing long-form video content across various demanding applications is expected to significantly reduce development overheads.
Brief written by urgent.news from Google Blog's own syndicated text. Machine-written — may contain errors; check the original before relying on it.
Also reported by 1 other outlet
- Introducing agentic video understanding with Gemini deepmind.google