
Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, expanding its native voice AI capabilities across more than 97 languages. The new models combine real-time conversation with visual understanding, tool use and deeper reasoning, pushing voice AI beyond simple assistants towards agents that can interact with software while continuing to speak with users.
Google is expanding the role of voice inside its Gemini ecosystem with the release of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two models designed specifically for real-time conversational applications.
Released on 15 September 2026 through the Gemini API and Google AI Studio, the models support more than 97 languages and are designed to maintain fluid dialogue while reasoning about requests and carrying out tasks. Google says Gemini 3.8 Live is focused on scalable, cost-efficient voice experiences, while the Extended Thinking version is built for more complex tasks requiring multi-step reasoning.
Voice AI that can act while it speaks
One of the more important changes is asynchronous function calling. This allows a Gemini-powered voice application to call APIs or other tools in the background while continuing its conversation with the user rather than pausing every time it needs external information.
That capability makes voice AI significantly more useful for agent-based applications. A customer service system could continue speaking with a user while checking an account or order status, while a workplace assistant could search internal systems, retrieve information and explain the result within the same conversation.
The models also support visual context, meaning voice agents can respond not only to what a person says but also to visual information supplied to the system. Google is effectively bringing speech, vision, reasoning and tool use into the same real-time interaction.
Why 97+ languages matters

Google’s Live API currently supports 97 languages, including Arabic, Chinese, Hindi, Portuguese, French and many languages that have historically received less attention from mainstream voice technologies. Its native audio models can also switch naturally between supported languages during a conversation.
The scale is particularly relevant because voice AI depends on more than translating text. Systems must interpret pronunciation, accents, rhythm, interruptions and changes between languages while still understanding the user’s intent.
Google has been investing heavily in this broader multilingual strategy. The company says its products and technologies now support interactions across more than 300 languages spoken by more than seven billion people, equivalent to around 86% of the world’s population.
Gemini 3.8 Live fits into that larger effort by making multilingual interaction part of an agentic AI system rather than treating language support as a separate translation layer.
Voice becomes an interface for AI agents
The bigger implication of the release is not simply that Gemini can speak more languages. It is that voice is becoming a way of controlling AI agents.
Earlier voice assistants mainly answered questions or carried out predefined commands. Systems such as Gemini 3.8 Live can increasingly interpret a request, reason through it, communicate with external tools and keep the user informed as the task is being completed.
Google is already bringing this approach into its productivity ecosystem. Earlier in September, the company introduced Gemini-powered conversational voice features across Gmail, Docs and Keep, allowing users to search email, draft documents and capture ideas through spoken interaction.
For businesses, the potential applications extend into customer service, healthcare, field work, retail, education and logistics, particularly in situations where workers need access to digital systems without constantly interacting with a screen.
However, the more capable voice agents become, the more important governance becomes as well. Systems that can access tools and perform actions need clear permissions, authentication and human oversight, particularly when they handle sensitive data or financial and operational decisions.
Gemini 3.8 Live therefore represents more than another improvement in synthetic speech. Google is building towards an environment in which speaking to an AI can trigger reasoning, visual interpretation and actions across digital systems.
The next stage of voice AI may not be defined by how human an assistant sounds, but by how reliably it can understand what people mean and turn conversation into controlled action.
Sources
- Google — Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe, 15 September 2026
https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/ blog.google - Google — Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, 15 September 2026
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/ blog.google - Google AI for Developers — Live API capabilities and supported languages
https://ai.google.dev/gemini-api/docs/live-api/capabilities Google AI for Developers - Google — AI for everyone in every language, 15 September 2026
https://blog.google/innovation-and-ai/technology/ai/ai-for-every-language/ blog.google - Google Workspace — Use your voice to get more done in Gmail, Docs and Keep, 3 September 2026
https://blog.google/products-and-platforms/products/workspace/voice-features-gmail-docs-keep/

Sara is a Software Engineering and Business student with a passion for astronomy, cultural studies, and human-centered storytelling. She explores the quiet intersections between science, identity, and imagination, reflecting on how space, art, and society shape the way we understand ourselves and the world around us. Her writing draws on curiosity and lived experience to bridge disciplines and spark dialogue across cultures.