Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking Bring More Natural Voice AI
Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-focused models designed to make conversations with AI more fluid while handling visual input, tool calls, and complex reasoning. The models are available through the Gemini API and Google AI Studio, with additional rollouts across Search, the Gemini app, Google Workspace, and enterprise products. as reported by Blog The release targets two different use cases. Gemini 3.8 Live is optimized for scale, cost efficiency, and responsive dialogue, while Gemini 3.8 Live Extended Thinking is designed for more demanding workflows that require multi-step reasoning.
Voice conversations that keep moving
Gemini 3.8 Live can process visual inputs in near real time, allowing it to use surrounding context during a conversation. Google’s examples include live employee onboarding, troubleshooting through Search Live, and playing chess while interpreting the board and responding conversationally.
The model can also detect and transition between 97 supported languages during a conversation. Meanwhile, tools and API calls can run in the background. That lets the model acknowledge a request and continue speaking while an external task is completed, rather than forcing the user to wait silently.
Google says Gemini 3.8 Live ranked second in the Speech Agent Arena and is built to remain cost-effective for developers and enterprises deploying voice agents at scale. The ai models story also surfaces in OpenAI’s Opaque Reasoning Technique Raises Alarm..., adding another angle.
Extended Thinking works while it speaks
Gemini 3.8 Live Extended Thinking is aimed at workflows where a quick response is not enough. Google says it can reason and speak simultaneously, using early verbal cues such as “Let me check that…” before providing a completed answer.
The model can also narrate progress during multi-step tasks. Demonstrations include turning rough sketches into React components, coordinating bookings through asynchronous function calls, and creating business plans and marketing toolkits through natural speech.
Google reports that the Extended Thinking model achieved an 82.6 score on Artificial Analysis’ Speech to Speech Quality Index. It also reported scores of 68.6% on τ-Voice for agentic task completion, 35.1% on Sierra’s τ-Voice-banking benchmark, and 97.7% on BigBench Audio. Google says the model holds the top overall position on the Speech to Speech Quality Index and maintains a competitive price compared with other frontier models.
On ServiceNow’s EVA-Bench, Google says both models push the benchmark’s Pareto Frontier for complex workflows by balancing conversational quality with accuracy. The company notes that the evaluation was run through the Live API on Gemini Enterprise Agent Platform. The ai models story also surfaces in Google AI Releases TimesFM 3 for..., adding another angle.
Availability for developers and users
Gemini 3.8 Live is rolling out through the Gemini API and Google AI Studio for developers. It is also in private preview for Gemini Enterprise, with availability planned for Gemini Enterprise for Customer Experience, and is rolling out in Search Live.
Gemini 3.8 Live Extended Thinking is likewise available through the Gemini API and Google AI Studio. Enterprise access is in private preview, with planned availability for Gemini Enterprise for Customer Experience and Google Workspace business customers. Users can access it through Gemini Live, while Google AI Pro and Ultra subscribers can use it in Workspace Docs; all Google AI subscribers can use it in Gmail and Keep.
Google is also supporting an ecosystem of voice-agent platforms, including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. The company says all audio generated by its AI products includes an imperceptible SynthID watermark intended to keep AI-generated audio detectable.
