Google propose Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are two new speech synthesis models that transform speech generation from a set of out-of-the-box options into a more flexible message creation tool. Models allow you to create your own voices, control the performance of individual lines, and generate expressive dialogue for games, podcasts, audiobooks, and other projects.
Image source: Google
Gemini 3.8 Flash TTS is designed for character creation and detailed control of their voices. Use text queries to define characters, accents, and voice characteristics in more than 100 languages and dialects. The library also contains more than 2,000 ready-made voices, including regional variants of Mexican Spanish, Quebec French, and Scottish English.
Additionally, Google has added the ability to copy sounds from 30-second recordings. This feature works with your own voices or voices that you’ve acquired rights to use. The system verifies the voice owner’s consent before creating a copy, and the resulting message is protected by a SynthID watermark and C2PA metadata.
Once a voice is selected, developers can control each voice individually using text instructions. Gemini is able to change the pace and character of a performance, using whispers and other vocal delivery elements, and adding reactions in the form of laughter, sighs, and brief comments.
Both models support long-form audio generation that maintains quality, natural rhythm and voice characteristics for hours, while dual-speaker mode allows you to create dialogue from a single script while maintaining separate voices and natural line order. Gemini 3.8 Flash-Lite TTS is primarily designed for large-scale projects, including dubbing, audio content creation, and voice AI agents.
According to Google, Gemini 3.8 Flash TTS took first place with a score of 71.4 on the Hume AI speech design benchmark and 60.8 on the accent modeling test. Flash TTS and Flash-Lite TTS also rank first and second respectively in the overall quality index metric (also from Hume AI). In blind user evaluations in Voice Arena, the model performed well in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

The new model is available starting today in Google AI Studio, which now has a separate interface for creating speech scripts. Gemini 3.8 Flash TTS and Flash-Lite TTS are also available through the Gemini API, and integrations have been announced for the Agora, LiveKit, Pipecat and Vercel platforms.
If you find an error, select it with your mouse and press CTRL+ENTER.









