Lilith Lilith.
Editorial illustration: Gemini 3.8 turns TTS from a voice menu into performance direction
Lilith illustration · editorial remix

Google has introduced two text-to-speech models. Gemini 3.8 Flash TTS targets character creation and detailed direction, while Flash-Lite TTS is designed for high-volume dubbing, audio production and voice agents. Both began rolling out on September 23 through the Gemini API and Google AI Studio.

A prompt creates the voice and each line gets direction

Flash can create a new voice from a description and direct each line through pacing, emotion, accent or conversational sounds. Google claims support for more than 100 languages and dialects and a library of over 2,000 production-ready voices. A custom profile can be saved for more consistent performance across a longer project.

The model can also replicate a voice from a 30-second recording. Google requires spoken consent from the owner of the reference voice for that feature. Gemini Notebook gets Flash, Google Vids gets Flash-Lite, and both are due to reach Gemini Enterprise through an API later.

Repeatable performance is the production advantage

For game, podcast or localization teams, precise direction matters more than another pleasant demo voice. Returning to the same character, changing one line and preserving an accent across scenes can shorten the cycle of generation, listening and correction.

Flash-Lite also reveals where Google expects volume: dubbing and voice agents. Splitting the models separates expensive creative control from workloads sensitive to price and latency. The announcement, however, gives no concrete price or latency figures.

Consent and watermarking do not govern a voice forever

Google says every output carries a SynthID watermark and also cites C2PA credentials for voice replication. Consent verification when a voice is created is a meaningful brake, but it does not by itself solve later license withdrawal, theft of a saved profile or use beyond the agreed project.

The benchmarks also remain largely part of the vendor story. Google reports 71.4 on Hume AI's Voice Design Benchmark and 60.8 for accent modeling. Buyers will care more about stability across long scripts and production workloads.

Rights, price and drift will decide adoption

Watch how easily a voice can be removed from a project, consent can be audited and one character can remain stable across hundreds of lines. Public Flash-Lite pricing, real Gemini API latency and the survival rate of SynthID after routine audio editing will matter just as much. Those operating numbers will separate a studio from an impressive playground.

Lilith's verdict

Google built a directing studio where an actor can emerge from 30 seconds of audio. It now has to prove that when the voice owner leaves the stage, the digital double actually stops performing.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗