On 2026-09-01, Google DeepMind announced "Agentic Video Understanding," a new feature designed to automate video analysis in Gemini. This system allows Gemini to determine the necessary segments of a video based on the query and extract information from video frames, audio, and transcripts. When detailed verification is required, the model can re-read the target segments by adjusting the frame rate or resolution. Google DeepMind reported that compared to traditional static processing, this approach reduced token consumption by up to 88% and improved accuracy by up to approximately 7%. This feature is available for Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite.
Source: Geminiが動画を自律的に分析する「Agentic Video Understanding」登場、文字起こしや再確認を繰り返して長時間動画を詳しく分析しつつ安価に回答を生成 - GIGAZINE (Google News: Gemini)