Google DeepMind launches two new Gemini TTS models for expressive voice creation

Google DeepMind introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, which generate custom character voices and allow line-by-line direction of pacing, emotion, and conversational sounds. The models are available across Google AI Studio, the Gemini API, and other products, and include watermarking for safety. They target applications such as audiobooks, podcasts, dubbing, and real-time voice agents.
The two new models serve distinct purposes within the Gemini Audio family. Flash TTS emphasizes creative control, letting users craft original character voices through natural language descriptions and direct line-by-line performances with adjustable acting cues and dialect shifts. Flash-Lite TTS prioritizes efficiency for high-volume tasks like dubbing and expressive voice agents.
Both models expand on the existing Gemini Audio lineup, which includes live translation and transcription tools. They offer access to over 2,000 production-ready voices across more than 100 languages and dialects, plus voice replication from just a 30-second sample. Deployment spans Google AI Studio, the Gemini API, Gemini Enterprise, Notebook, and Google Vids, with built-in watermarking for audio security.
These tools could significantly lower production barriers for independent creators, small studios, and developers seeking professional-grade voice work without hiring voice talent. Audiobook producers, podcasters, and game developers may find new creative flexibility, while voice actors could face increased competition from synthetic alternatives. The watermarking feature suggests an effort to address authenticity concerns, though broader questions about voice ownership and consent in replication scenarios may emerge as adoption grows.