Google launches Gemini 3.8 Flash TTS with prompt-built voices and 30-second cloning
- Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Sep 23, 2026, calling them its most expressive audio generation models yet, for use in Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
- Gemini 3.8 Flash TTS creates entirely new character voices from natural language prompts and directs each line with control over acting cues, pacing, dialect shifts, and backchanneling, while Flash-Lite TTS is tuned for high-volume dubbing and voice agents.
- Voice replication rebuilds a vocal profile from a 30-second audio sample, guarded by consent verification, SynthID watermarking, and C2PA credentials on the generated audio.
- Commenters estimate audiobook generation costs roughly $5 to $10 per 10 hours with few-shot prompting: $0.81/hour standard or $0.41 batch for 3.8 Flash TTS, $0.54 or $0.27 for Flash-Lite.
- Early users report prompt adherence is weak in the demos: the sample labeled "super tinny monotone robotic voice" is neither tinny nor monotone, and one tester hit errors in Voice Design and found voice replication blocked in their region.
Hacker News opinions
I built KeenLore, a locally hosted emotive audiobook creator. Gemma 4 does the prose analysis and Qwen3 TTS Voice Design makes the voice samples; quotation attribution hits 97.2% (485/499) on my novel, all on an 8GB T1000 and a Ryzen 5 7600.
Cool tech and I listen to hundreds of hours of TTS, but my brain already fills in character voices from the text on the page. Multiple voices are fun, just not that critical to me.
I've done a similar project all year. Try Fish Audio or Higgs instead of Qwen3, both give much better prosody and are far easier to listen to over long runs.
That voice replication blurb is telling. Google cloning from a 30-second sample with consent checks and SynthID watermarking basically means they gave up holding back since everyone else ships cloning now.
They probably do what GPT-Live does: the voice profile has to record a line saying it consents to synthetic samples. Or local cloning is already good enough that Google granting it adds no uniquely liable ability.
Last night I fine-tuned QwenTTS 1.7B on clean recordings and got a robo-me that sounds absurdly good in a few hours. My family was shocked. The cat is out of the bag for sure.
This has worked for 1-2 years already. There are multiple open models that clone voices well, so Google's release is not new capability.
The demo labeled "super tinny monotone robotic voice" is neither tinny nor monotone. Sounds nothing like 90s TTS or movie robots, more like hype DJs.
None of these examples are what the prompt asked for. Same as image models: once the shock wears off you realize the output isn't what you wanted.
I direct my own Star Trek fanfic and GPT-Live won't give me distinct voices or the expressiveness I imagine. Line-by-line direction in Gemini 3.8 is what I need, though I still can't tell which of the 5,286 Gemini products to use, and I worry what training happens to my text files.
For audio drama try alexandria-audiobook, or Qwen3-TTS where you just describe the voice. I run pdf-narrator over Kokoro on an 8GB M1 MacBook Air and make my own audiobooks for free, favorite voice is am_michael.
How much to batch generate an audiobook? I use the ElevenLabs app free right now and would rather just export audio files.
Roughly $5-10 for 10 hours if you few-shot it. Per hour: 3.8 Flash TTS $0.81 standard, $0.41 batch; Flash-Lite $0.54 standard, $0.27 batch. Great price at least until December 31.
This should power the Google Play Books app, its voice system is ancient. It's the single best place to apply AI and it's still untouched.
The ratio of new voice models I see on Hacker News to the ones actually deployed in any product I use is approximately infinity.
I've never heard a TTS do a convincing British accent. They all sound like Americans doing their best fake one.
Voice actors are safe for now. Technically impressive, but the results are not very close to the prompt, some core part of the request gets ignored in basically every demo.
I listened to a bunch: male voices are believable, all the female voices sound the same and artificial, weirdly like Toy Story. Training data bias?
Rollout is sloppy. Voice Design errors on me in AI Studio, voice replication is blocked in my region, no neutral gender voices in English, the Gaming use case is empty, and no pricing page anywhere. Available voices just sound like generic Gemini voices.