Lilith Lilith.
Editorial illustration: Gemini launches live audio interface
Lilith illustration · editorial remix

Native voice competence without external translation

Google has released two new speech models: Gemini 3.8 Live and an Extended Thinking variant. This is not a standard ASR (speech recognition), LLM (text), and TTS (synthesis) pipeline, but models with native audio processing. They handle over 70 languages and can retain the intonation, pacing, and pitch of the original speaker.

Who controls the default communication channel

The battle between OpenAI and Google is shifting from the text box to fluid, real-time dialogue. For enterprises, this opens the possibility of building smoother interactions into customer support and interpretation services. The key advantage is lower latency, as the text conversion step is eliminated.

Multilingualism hits local nuances

On paper, the models handle hundreds of language pairs without switching settings. However, the real limit won't just be word recognition, but the ability to maintain context and idioms in edge languages. Quality will likely differ significantly between global English and smaller European markets.

Latency under infrastructure pressure will decide

The most interesting metric will be the stability of response times once the system faces typical operational load. The proof that the architecture works will be the absence of noticeable pauses when translating complex sentences.

Lilith's verdict

We are watching a race to see who will become the default interpreter for global business and cut off old call centers from revenue.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗