The Daily Commit · Section Edition Front Page PHP AI Dev EN DE FR ES

TheModelDesk

September 15, 2026
models, agents & local inference

Releases

Google expands Gemini 3.8 with two live voice models

Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026. The models bring near-real-time voice interaction, visual context, background tool execution, and 97-language switching to the Gemini API. The extended model adds deeper reasoning while speaking. Google is rolling both out through developer tools, enterprise previews, Search, Gemini, and selected Workspace services.

Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026. Both models target voice agents that can reason during near-real-time dialogue. The announcement is signed by Tom Ouyang, Principal Engineer, and Malini Jaganathan, Member of Technical Staff on behalf of the Gemini Audio Team.

Gemini 3.8 Live is designed for scale and cost efficiency. It combines conversational intelligence, fluid dialogue, and visual grounding. The model placed second in Artificial Analysis’ Speech Agent Arena. On ServiceNow’s EVA-Bench, Google says both models push the Pareto frontier for complex workflows by balancing accuracy with conversational quality. The test used the Live API on Gemini Enterprise Agent Platform.

Gemini 3.8 Live Extended Thinking took the top overall position in Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. It reached 68.6% on τ-Voice for agentic task completion and 35.1% on Sierra’s τ-Voice-banking benchmark. Its Big Bench Audio score is 97.7%. Google describes the model as suitable for complex enterprise tasks and says its price remains competitive with other frontier models.

Gemini 3.8 Live processes visual input in near real time and can switch automatically among 97 supported languages during a conversation. It runs tools and API calls in the background while the dialogue continues. Demonstrations show the model helping with live employee onboarding and playing chess using visual context. Extended Thinking can reason and speak simultaneously, acknowledge a request with an early cue such as “Let me check that…”, and narrate progress during multi-step background work. Other demonstrations show sketches becoming functional React components, multi-step bookings using asynchronous function calls, and business plans with custom marketing toolkits created through speech.

Google says the models improve voice interaction across the Gemini app, Google Workspace, and Search. Gemini 3.8 Live Extended Thinking is available in Docs Live, Gmail Live, and Keep Live. Gemini 3.8 Live provides step-by-step troubleshooting through Search Live.

The Gemini Live API is supported by Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms handle real-time media streaming infrastructure for developers building voice interfaces. Google also names Salesforce, Genspark, and Lumeris as partners interested in the models’ latency, conversational flow, and tool-calling capabilities.

Audio generated by Google’s AI products carries an imperceptible SynthID watermark. Google says the watermark keeps generated audio detectable and supports efforts to limit misinformation. Further safety details are provided in the Gemini 3.8 Audio model card.

Both models begin rolling out on September 15, 2026. Developers can use them through the Gemini API and Google AI Studio. Enterprise access starts in private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience announced for a later release. Gemini 3.8 Live is also rolling out in Search Live. Extended Thinking is available in Gemini Live, in Workspace Docs for Google AI Pro and Ultra subscribers, and in Gmail and Keep for all Google AI subscribers. It is also coming to Google Workspace business customers.

Read the original source ↗

Rate this article: 0

Readers’ Forum

No contributions yet — open the debate.

← The Model Desk — Page C1

Models, agents & local inference · The Daily Commit · Screen edition · Imprint · Privacy Policy