Google has introduced Gemini 3.8 Live with Live Avatar, adding near real-time visual presence to its conversational AI models.
The system combines speech with streaming video to create dynamic avatars with precise lip-syncing, natural expressions, and fluid turn-taking. It processes visual and audio inputs together and supports asynchronous tool calls, allowing agents to retrieve information while conversations continue.
Live Avatar also supports speech-to-speech interactions across 97 languages and lets enterprises create custom avatars from reference images through allowlisted access. Audio and video outputs include SynthID watermarking for transparency. Gemini 3.8 Live with Live Avatar is available through Gemini Enterprise.




