Google published releases on September 15, 2026, for two distinct voice models: Gemini 3.8 Live, designed for cost efficiency, visual grounding, and conversational dialogue, and Gemini 3.8 Live Extended Thinking, targeted at multi-step reasoning and complex tasks. Developers can access both variants via the Gemini API, Google AI Studio, and Google Cloud's Vertex AI, while access across Google's first-party products splits depending on the workflow.

Conversational Live Flow vs. Extended Thinking

Both releases build on Gemini 3 Pro, according to Google's model card, accepting multimodal context—audio, images, video, and text—up to 128,000 tokens, with outputs reaching up to 64,000 tokens in audio and text. The operational split between the two variants addresses different agent interaction patterns:

  • Gemini 3.8 Live prioritizes fluid, near-real-time spoken dialogue and visual input processing. Google claims it supports transitions between 97 languages mid-conversation and allows tool or API execution in the background while the audio stream remains active.
  • Gemini 3.8 Live Extended Thinking focuses on higher-complexity tasks. Google states the model can speak and reason simultaneously, offering early verbal cues and narrating progress while executing multi-step operations in the background.

This division shifts voice agents away from simple turn-by-turn question-answering systems toward interfaces capable of continuing an interactive voice session while slower external tool executions or analytical processes finish.

Distribution Channels and Workspace Splits

While both variants share developer access points, their integration into Google's application ecosystem differs:

ModelDeveloper ChannelsApplication Surfaces
Gemini 3.8 LiveGemini API, Google AI Studio, Google Cloud / Vertex AIGemini App, Google Search Live
Gemini 3.8 Live Extended ThinkingGemini API, Google AI Studio, Google Cloud / Vertex AIGemini App, Google Workspace (Gmail, Docs, Keep)

The model card specifies these specific surfaces rather than a blanket global rollout. Neither pricing schedules nor geographic rollout schedules were provided in the release material.

Benchmark Claims and Model Limitations

Google reported evaluation figures for the models, but these are company-reported results rather than independently audited assessments. Google reported an 82.6 on Artificial Analysis's Speech to Speech Quality Index, 68.6% on tau-Voice, and 35.1% on Sierra's tau-Voice-banking benchmark. On Big Bench Audio, Google reported 97.7% for Live Extended Thinking.

The model card also outlines material operational boundaries. Both models have a knowledge cutoff of January 2025 and are subject to hallucinations, occasional latency spikes, or timeouts during execution. Regarding safety, Google stated its frontier-safety assessment found no meaningful new capabilities or material performance gains compared to Gemini 3.7 Flash on the evaluated dimensions, and Google indicated it does not expect the models to reach critical capability risk tiers.