Overview
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are two near-real-time voice conversation models released by Google DeepMind on September 15, 2026, which the company calls its most advanced real-time conversational models to date. The two models launched the same day and share the same foundation, differing in the scenarios they target: Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence, fluent interaction and visual grounding; Gemini 3.8 Live Extended Thinking targets highly complex tasks, with stronger multi-step reasoning that lets it reason and speak at the same time.
The interaction design of Extended Thinking is the most interesting change in this generation. It confirms requests with advance verbal cues like "Let me check that…", then keeps narrating its progress by voice while pushing multi-step tasks forward in the background, turning the wait into a perceptible conversation instead of leaving the call in silence. Both models can execute tool calls and API requests in the background while continuing to talk with the user, and both let users interrupt at any time and change their mind on the spot.
On the capability side, both models can automatically detect and switch languages mid-conversation, covering 97 languages; all generated audio carries an imperceptible SynthID watermark. Developers access them through the Gemini Live API, with a /live entry point in Google AI Studio. On the enterprise side they are in private preview via Gemini Enterprise, while on the consumer side they have already arrived in Search Live, Gemini Live, and Workspace's Docs, Gmail and Keep.
Key Features
- Two variants with clearly divided roles: Gemini 3.8 Live focuses on scale and cost efficiency, handling fast, fluent real-time conversation; Extended Thinking targets highly complex tasks, using stronger reasoning for multi-step task execution and background processing.
- Reasoning while speaking: Extended Thinking outputs speech in sync with its reasoning, confirming requests with advance verbal cues like "Let me check that…" and narrating the progress of multi-step background tasks in real time, so users don't have to wait in silence.
- Background tool calls and asynchronous long tasks: Executes tool calls, API requests and asynchronous function calls in the background while keeping natural conversation uninterrupted, ideal for voice scenarios that require looking up data, placing orders or walking through a process.
- Visual grounding: Processes visual input in near real time, bringing what's in the user's camera or on their screen into the conversation context to support real-time guidance tasks where you ask questions while watching.
- Mid-conversation switching across 97 languages: Automatically detects language and supports switching languages within the same conversation, with no need to restart the session.
- Interruptibility and parallel reasoning: Lets users interrupt at any time and correct course; this generation brings a major upgrade in intelligence and parallel reasoning, handling the main conversation and background tasks simultaneously.
- SynthID audio watermarking: All generated audio is woven with an imperceptible watermark, making AI-generated content easier to identify and reducing the risk of forgery.
Use Cases
- Voice customer service and agent assistance: look up information and run business processes while on the call, without putting users on hold
- Real-time troubleshooting: combine visual input and voice to guide users step by step through a fix
- Employee onboarding for enterprises: answer real-time questions during a call and tie in visual context
- Multi-step bookings and transactions: call functions asynchronously in the background while keeping the conversation flowing naturally
- Voice products for global users: cover multilingual scenarios with real-time switching across 97 languages
- Voice work inside Google Workspace: real-time voice collaboration in Docs, Gmail and Keep
Pros
- Officially positioned as the most advanced real-time conversational model to date, ranking first on the Speech to Speech Quality Index at 82.6
- Extended Thinking runs reasoning and expression in parallel, so the conversation never goes cold during long tasks
- Background tool calls and asynchronous tasks don't block the current session
- Supports 97 languages and switches automatically mid-call
- Visual grounding lets the voice assistant read the scene in front of it
- There is a broad ecosystem of integrators, with platforms including LiveKit, Agora, Pipecat, LangChain, Vercel and Vision Agents all supporting the Gemini Live API
Pricing
Google has not yet disclosed specific pricing, describing it officially as highly competitive relative to other frontier models and emphasizing the cost efficiency of 3.8 Live. Developers can try it through the Gemini API and Google AI Studio (the /live entry point); Gemini Enterprise is in private preview; Search Live is available directly to general users, Extended Thinking is available in Workspace Docs for Google AI Pro and Ultra subscribers, and all Google AI subscribers can use it in Gmail and Keep. Refer to the official page for actual pricing.
Summary
The Gemini 3.8 Live family addresses the two things users complain about most in voice assistants: long waits and not daring to interrupt. The standard version makes real-time conversation cheap and reliable, while Extended Thinking lets complex tasks run to completion within a call. Among the official results, Extended Thinking ranks first on the Artificial Analysis Speech to Speech Quality Index with 82.6, plus 68.6% on τ-Voice and 97.7% on Big Bench Audio, with voice agent task completion the focus of this upgrade cycle. Products handling voice customer service, real-time guidance or multi-step transactions should prioritize testing this generation of the Live API; if you only need text conversation, a general-purpose model like Gemini 3.8 Flash remains more straightforward.
Version History
- Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking officially launch (2026-09-15): Google DeepMind released two near-real-time voice conversation models: Gemini 3.8 Live, aimed at scale and cost efficiency, and Gemini 3.8 Live Extended Thinking, aimed at highly complex tasks with support for reasoning while speaking. Extended Thinking scored 82.6 on the Artificial Analysis Speech to Speech Quality Index to rank first, along with 68.6% on τ-Voice, 35.1% on τ-Voice-banking and 97.7% on Big Bench Audio; the standard version placed second on the Speech Agent Arena. Both models support mid-conversation switching across 97 languages, user interruptions, background tool calls and asynchronous long-running tasks, with all audio carrying a SynthID watermark. Developers can access them through the Gemini API and AI Studio, the enterprise side is in private preview as Gemini Enterprise, and the consumer side has launched Search Live, Gemini Live and Workspace's Docs, Gmail and Keep.