FR
live
tag

#speech-to-text

Google closes the multimodal loop with Gemini 3.5 Transcribe and the GA release of Omni 1.1 Flash for video

On 26 August 2026, Google made Gemini 3.5 Transcribe generally available, two dedicated speech-to-text models with diarization and custom vocabulary, and on 27 August it shipped Gemini Omni 1.1 Flash, its conversational video generation model with interpolation and 4K output. Transcription is no longer a feature of the generalist model — it is a standalone product. Here is what that changes for teams that transcribe or produce video.

Type at least two characters.

navigate open esc dismiss