Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS. These are two new models for speech generation. Flash TTS can create new voices from text descriptions. Both models support more than 100 languages. They let users add stage directions to individual lines of dialogue.
Gemini 3.8 Flash TTS targets creative projects such as game characters, audiobooks, and podcasts. Gemini 3.8 Flash-Lite TTS focuses on low-cost speech generation at scale for dubbing, audio content, and voice agents.
Gemini 3.8 Flash TTS lets users design voices from scratch. A text prompt can define a voice role, accent, and vocal traits across various languages and dialects. Google offers a library of more than 2,000 preset voices for users who prefer not to start from zero. This library includes regional variants like Mexican Spanish, Quebec French, and Scottish English.
A voice cloning feature builds a voice profile from a 30-second audio sample. The person whose voice is cloned must record a spoken statement of consent. The voice in that recording must match the sample. Every clip generated by the Gemini audio models carries an inaudible SynthID watermark to help detect AI-generated speech.
Google also announced Voice Remixing. This feature will let users adjust the timbre, pitch, tempo, and accent of library voices. It is not available yet.
Both models let users write directions for each line or have the model interpret script cues independently. The models can generate hours of audio with minimal speaker drift. This means the voice barely changes over time.
A two-voice mode generates dialogue from a single script while keeping the voices distinct. Users can script laughter, sighs, and sounds like mhm to place reactions and pauses precisely where they want them.
Google is rolling out both models through the Gemini API and Google AI Studio. Flash TTS is also available in Gemini Notebook. Flash-Lite TTS is available in Google Vids. Access through the Gemini Enterprise API will follow soon for both models.
Developer platforms including Agora, LiveKit, Pipecat, and Vercel already support integration through the Gemini API. Google has not listed regional endpoints for the new models yet. Earlier TTS models offered EU data processing.
Data from the free tier is used to improve products. Data from the paid tier is not. Paid pricing is listed in US dollars per million tokens. Text tokens are billed for input. Audio tokens are billed for output.
One second of generated audio equals 25 audio tokens. An hour equals 90,000 tokens. An hour of audio output costs 81 cents with Flash TTS and 54 cents with Flash-Lite TTS through the end of 2026. Costs rise to $1.62 and $1.08 on January 1, 2027. Text input is billed separately.



