Google has launched the general availability of Gemini 3.8 Flash, a new model designed for extended coding sessions, autonomous agents, and complex workflows. It supports a context window of up to 1 million tokens and output of up to 64,000 tokens.
Model ReleasesGoogleGemini 3.8 Flash
Google Launches General Availability of Gemini 3.8 Flash, Optimized for Agents
This article is a translation. Read the Japanese original
Rather than simply answering single questions, this model features "multi-step planning," enabling the AI to create a plan to achieve a goal, iteratively call tools, and verify and correct the results to complete a series of tasks. Its application as an AI agent is emphasized, as it has been adopted as the default model for Google Managed Agents.
Additionally, the model includes a feature that allows users to select the reasoning level from three stages: "low," "medium," and "high." This enables resource allocation based on the application, such as "low" for chats requiring low latency and "high" for deep inference or mathematical tasks. The default setting is "medium."
Regarding API usage, there are specification changes, such as the replacement of the traditional thinking_budget with thinking_level. As for pricing, a special price has been set until 2026-12-31, with a planned transition to standard pricing on 2027-01-01.
Sources
- Is the goal of "Gemini 3.8 Flash" not "intelligence"? Three key points from the announcement (ITmedia AI+, 2026-09-08)