Intelligent transcription with Gemini 3.5 Transcribe

Google DeepMind introduces Gemini 3.5 Transcribe, a speech-to-text model available via the Gemini API for developers. It supports real-time streaming with sub-second latency and pre-recorded audio processing with speaker attribution. The model handles noise, jargon, and disfluency cleanup, enabling voice agents and captioning tools.

Cover image for Intelligent transcription with Gemini 3.5 Transcribe