English

Model ReleasesGoogleGemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS

Google Announces Gemini 3.8 Flash TTS, Enabling Voice Creation from Scratch Using Natural Language

This article is a translation. Read the Japanese original

On September 23 (local time), Google announced its Text-to-Speech (TTS) models, "Gemini 3.8 Flash TTS" and "Gemini 3.8 Flash-Lite TTS."

Flash TTS is designed for character voice creation and detailed performance direction, making it suitable for applications such as games, audiobooks, and podcasts. On the other hand, the more affordable Flash-Lite TTS focuses on cost efficiency, making it ideal for high-volume dubbing, voice agents, and similar use cases.

A major feature of these new models is the ability to generate entirely new voices from scratch through natural language instructions, moving beyond the conventional method of selecting from preset voices. Once created by specifying roles, accents, or vocal qualities, these voices can be saved, allowing for consistent use across long-term projects. Additionally, the models include a "voice replication" feature that can reproduce a voice from a 30-second audio sample. Please note that using replication requires the speaker's verbal consent and the submission of a recording.

In terms of performance, users can control speaking pace, emotion, non-verbal sounds such as laughter or sighs, and the insertion of interjections by providing stage-direction-like instructions for every 1 lines. Generated audio will include "SynthID," a digital watermarking technology designed to detect AI-generated content.

Flash TTS will be available through the Gemini API and Google AI Studio, and general users can access it via Gemini Notebook. Flash-Lite TTS will be available through similar channels, as well as through Google Vids for general users.

Sources

  1. ゼロから声を作れる音声生成モデル「Gemini 3.8 Flash TTS」公開 「Gemini Notebook」でも利用可能に (ITmedia AI+、2026-09-24)