Lilith Lilith.
⌕
Editorial illustration: Gemini 3.8 builds a voice from 30 seconds, but leaves Europe outside
Lilith illustration · editorial remix

Google has released two text-to-speech models. Gemini 3.8 Flash TTS targets creative direction, acting and long dialogue, while Flash-Lite is designed for cheaper high-volume workloads. The catalogue contains more than 2,000 ready-made voices, and both models can stage two speakers from one script.

Thirty seconds of audio can become a persistent voice

Flash TTS can design a new voice from a text description or replicate an existing one from a 30-second recording. Google requires a consent recording from the voice owner and marks generated audio with SynthID and C2PA credentials. Custom voices can be saved for later projects, turning the feature into production infrastructure rather than a disposable effect.

Simon Willison built a browser playground on top of the API. His 1 minute 18 second dialogue took about 20 seconds to generate and cost 2.74 cents. That operating datapoint is more useful than a benchmark rank on its own.

Direction, not synthesis alone, is the production advantage

For podcast, dubbing and voice-agent teams, the useful layer is control over individual lines, pacing, accents and vocal reactions. One script can keep two characters and their voice instructions separate. Flash prioritizes fidelity and character work, while Flash-Lite prioritizes throughput and cost.

Portability is part of the shift. A stored voice identifier can reduce drift across episodes and shorten prompts, allowing a voice to behave like a managed production asset.

The most valuable feature comes with a geographic brake

Voice replication in Google AI Studio is unavailable in the EEA, UK, Switzerland, India, Illinois and Texas. European teams can test the models and design synthetic voices, but they cannot replicate their own voice in AI Studio. Consent checks, watermarking and C2PA reduce misuse risk, yet they do not by themselves settle rights management after a performer leaves a project.

Consent management after creation will decide the product

The telling signal will be whether Google provides usage audits, consent withdrawal and reliable deletion of a voice profile across products. Long-form stability over hours, rather than a short demo, is the second test. Production will reveal whether 2,000 voices accelerate delivery or merely multiply the options awaiting approval.

Lilith's verdict

Google has opened a studio with 2,000 voices, but the European performer still cannot reach the cloning booth. Gemini TTS will be judged by whether a voice can be withdrawn after the final take as reliably as it can be created.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗