Google has introduced new Gemini Audio models for developers building real-time voice applications. Gemini 3.8 Live supports speech-to-speech conversations while executing tasks through asynchronous function calls, processing visual context, and maintaining dialogue across more than 97 languages.
Gemini 3.8 Live Extended Thinking adds configurable reasoning for complex, multi-step voice tasks. Google also highlighted Gemini 3.5 Transcribe, a dedicated speech-to-text model supporting more than 85 languages, automatic code-switching, custom vocabulary, and structured transcripts.
The Live models are available through the Gemini Live API, while developers can access the broader audio suite through Gemini API and Google AI Studio.




