Google says the feature can run across web, mobile and interactive kiosks. Enterprise customers can use US and EU endpoints, provisioned throughput, and Google Cloud's enterprise compliance and data-governance controls.

The Live Avatar update builds on Gemini 3.8 Live, which Google introduced the previous week. Google describes the feature as processing audio and visual input together and generating near-real-time video alongside speech. The company also says it supports asynchronous tool calls, so an agent can fetch information in the background while continuing a conversation, and native speech-to-speech synchronization across 97 languages. The announcement does not provide independent performance measurements for those claims.

Custom avatar creation is available only through an enterprise allowlisting and verification process. Google says generated audio and video include imperceptible SynthID watermarks. Gemini 3.8 Live Extended Thinking is still in private preview, according to Google Cloud.

The release gives enterprise teams a generally available product and API path for conversational agents that combine live audio, video and tools. Teams evaluating it will need to assess quality, latency and operating costs against their own workloads; the supplied announcements do not report independent benchmarks or pricing.