Google announced the speech-to-text model Gemini 3.5 Transcribe on August 26, 20261. It arrives in public preview through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, split into two APIs by use case: gemini-3.5-transcribe-live on the Live API, which streams bidirectionally with sub-second latency for interactive voice apps, and gemini-3.5-transcribe on the Interactions API, which transcribes recorded audio with speaker attribution and word-level timestamps1.
On accuracy, Google cites measurements by Artificial Analysis: an average word error rate of 4.0% for streaming and 2.6% for non-streaming use cases1. The model automatically detects and transcribes over 85 languages and handles self-corrections and filler-word removal, though speaker attribution in pre-recorded audio covers up to three speakers, and Google notes that support for 3+ speakers is experimental1. It already runs behind Gboard’s Rambler on Android and the Gemini app on macOS; Chrome is listed as coming soon, and the post gives no pricing1.
Sources
- Intelligent transcription with Gemini 3.5 Transcribe - Google official blog (August 26, 2026)