Google DeepMind Launches Two New Voice-First AI Models for Real-Time Dialogue

Google DeepMind has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which are designed to handle complex reasoning, real-time visual context, and background task execution during voice conversations. The models are available through the Gemini API, Google Workspace, and the Gemini app, aiming to make voice interactions more natural and fluid. They support parallel reasoning and can manage tools while the user continues chatting.
The two models serve distinct purposes within the same voice-first framework. Gemini 3.8 Live prioritizes speed and cost efficiency for everyday conversational use, while the Extended Thinking variant targets enterprise-grade tasks requiring multi-step reasoning. Both support parallel processing, allowing the AI to manage tools and background operations without halting the user's speech flow.
Benchmark results highlight the Extended Thinking model's strengths: it ranked first on Artificial Analysis' Speech to Speech Quality Index with an 82.6 score, achieved 68.6% on the τ-Voice agentic benchmark, and scored 97.7% on Big Bench Audio. The standard Live model placed second in the Speech Agent Arena while maintaining a competitive price point for developers.
These models could meaningfully shift how people interact with AI assistants in daily life, moving voice interfaces from simple command-response systems toward genuine collaborative dialogue. Professionals juggling complex tasks—analysts, developers, customer service agents—may find hands-free multitasking increasingly viable, potentially altering workplace workflows. However, the emphasis on natural, fluid conversation also raises questions about user reliance on AI for reasoning tasks, which could affect critical thinking habits over time. The competitive pricing suggests broad adoption is feasible, but the societal impact will depend on how transparently these systems communicate their limitations during extended interactions.