English

Model ReleasesGoogleGemini 3.5 Transcribe

Google Releases Gemini 3.5 Transcribe Speech-to-Text Model

This article is a translation. Read the Japanese original

Gemini 3.5 Transcribe is a Speech-to-Text model that converts audio directly into formatted text. It is specifically optimized for handling background noise, technical terminology, and filler words.

The model is provided through two types of APIs. One is a model strong in intent interpretation and custom vocabulary, while the other supports streaming transcription and language switching.

Performance has improved significantly. According to measurements by Artificial Analysis, the time to final transcription has improved by 70% compared to the previous generation, Chirp 3.

Multilingual performance has also been enhanced. Google stated that the model achieved a Word Error Rate (WER) of 5.50% in streaming mode and 5.04% in non-streaming mode on the FLEURS benchmark.

This model is already being utilized in Android's Rambler and the Gemini app for macOS. Additionally, it is possible to build high-performance voice interfaces through platforms such as LangChain and Vercel [HN Search (backfill)].


Source: Gemini-3.5-Transcribe (HN 362pt, 127 comments) (HN Search (backfill), 2026-08-28)