Google launches Gemini 3.5 Live Translate: speech-to-speech model for 70+ languages
On June 9, 2026, Google launched Gemini 3.5 Live Translate — a model that translates speech to speech in real time across 70+ languages. Audio is generated continuously, staying several seconds behind the speaker. The feature is available via Gemini Live API, Google Meet, and Google Translate app.
AI-processed from MarkTechPost; edited by Hamidun News
Google Launched Gemini 3.5 Live Translate: Speech-to-Speech Model for 70+ Languages
Google released Gemini 3.5 Live Translate — a streaming model for simultaneous real-time speech translation supporting over 70 languages. According to MarkTechPost on June 9, 2026, the new model translates speech directly to speech, generating audio as a continuous stream with just a few seconds lag behind the speaker. The development is available through the Gemini Live API, and is also built into Google Meet and the Google Translate application.
How Streaming Translation Works
Unlike classical machine translation systems that first recognize speech completely, translate the text, and only then synthesize audio, Gemini 3.5 Live Translate operates on a streaming principle: the model processes incoming sound in fragments and begins generating translated speech almost immediately, without waiting for the speaker to finish their phrase. This is precisely why the delay is reduced to just a few seconds — a latency comparable to professional simultaneous interpretation at a conference, rather than the familiar pause in translator apps that process the entire recording.
Key facts about the launch:
- The model is named Gemini 3.5 Live Translate and belongs to Google's Gemini family.
- Supports speech translation into more than 70 languages.
- Works in streaming speech-to-speech mode with just a few seconds lag from the original.
- Available through Gemini Live API for third-party developers.
- Integrated into Google Meet and the Google Translate application.
Where the Model Will Appear
Google is embedding Gemini 3.5 Live Translate into several of its products at once. In Google Meet, the model will be able to provide conversation translation between participants in video calls in different languages with virtually no delay, which closes a long-standing gap in video communication services — until now most of them relied on subtitle translation rather than live voice translation. In the Google Translate application, the new model replaces or complements previous voice translation mechanisms, making the conversational mode of the application more natural for travelers and business negotiations.
Separately, the model is open to developers through the Gemini Live API — a programming interface that allows embedding real-time functionality (audio, video, text) into third-party applications. This means that speech-to-speech translation will be usable not only by Google's own products but by any companies creating call centers, educational platforms, customer support assistants, or systems for international events.
A Long Path to Live Speech Translation
The idea of simultaneous machine speech translation is not new: voice translator earbuds and apps with "phrasebook" functionality have existed for years, but almost all of them suffered from the same problem — a noticeable pause between the remark and the translation, which turned natural dialogue into an exchange of monologues with waiting. Each link in the classical pipeline — speech recognition, machine text translation, voice synthesis — added its own delay and its own source of errors, and mismatches between stages could distort intonation, pauses, and even meaning. The streaming architecture that combines all these steps into a single model is the answer to precisely this problem, and the fact that Google is releasing it immediately into mass products like Meet and Translate, rather than keeping it in experimental status, speaks to the company's readiness to scale the technology to millions of users.
What This Changes for Developers and Business
The availability of a streaming translation model accessible via API lowers the barrier to entry for companies that previously either had to hire a team of translators or integrate disparate services for speech recognition, text translation, and voice synthesis separately — with corresponding delays at each stage. A unified speech-to-speech model simplifies the architecture of such products and potentially reduces cumulative latency and service costs.
For end users, the main effect is bringing machine translation closer to the experience of live simultaneous interpretation in everyday scenarios: video calls, trips, negotiations. Given that Google is simultaneously developing Gemini as an ecosystem of models and embedding new capabilities into its mass products — Meet, Translate, search and office services — streaming speech translation could quite quickly become a standard feature rather than a niche capability available only through specialized simultaneous translation services.
The question of translation quality in complex scenarios remains open — with overlapping remarks from multiple speakers, strong background noise, or highly specialized terminology, where even professional simultaneous interpreters make mistakes. The few seconds of lag Google mentions is a compromise between speed and accuracy: the less context the model manages to accumulate before generating a translation, the higher the risk of inaccuracies on long and syntactically complex phrases. Nevertheless, the very fact that such functionality is now built into everyday tools like Google Meet, rather than requiring separate specialized equipment, significantly lowers the barrier to international communication in business and everyday life.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.