Gemini 3.8 text-to-speech lets you design a voice from a description, or copy yours from 30 seconds

20 hours ago

Google's Gemini 3.8 Flash TTS and Flash-Lite TTS became generally available on September 22, 2026. Describe a voice in words or copy one from a 30-second sample with its owner's recorded consent, direct each line, and hear them in Gemini Notebook and Google Vids. Plus what changes for people on the older preview model. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ https://ai.google.dev/gemini-api/docs/changelog https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts https://ai.google.dev/gemini-api/docs/pricing

Ask

Ask about this presentation

Answers are generated from this presentation.

Chapters

  1. 0:00Gemini 3.8 Flash TTS and Flash-Lite TTS
  2. 0:23Two models, two jobs
  3. 0:48Voices of your own
  4. 0:50Describe a voice and Gemini makes it
  5. 1:17Copy a voice from 30 seconds, with the owner's consent
  6. 1:43Direct the performance line by line
  7. 2:04Who it reaches
  8. 2:06Where you can use them
  9. 2:26On the older preview model? Check your scripts
  10. 2:52Where to read more
Show transcript

Gemini 3.8 Flash TTS and Flash-Lite TTS

010203What shipped
Gemini API release notes · Sep 22, 2026 · blog.google · Gemini 3.8 text-to-speech
Gemini

Gemini 3.8 speech: design a voice, or copy yours

Generally available September 22, 2026Gemini API and AI Studio

Google has released two new Gemini models that turn text into speech: Gemini 3.8 Flash TTS, and Gemini 3.8 Flash Lite TTS. They became generally available in the Gemini API on September 22nd. The news is that you're no longer choosing from a fixed set of voices. You can describe a new one, or copy a real one.

Two models, two jobs

Gemini 3.8 TTS
010203What shipped
ai.google.dev · Gemini 3.8 TTS model pages · ai.google.dev · Gemini API pricing
ModelBuilt forLanguages
Flash TTSaudiobooks, characters, dialogue130
Flash-Lite TTSdubbing, voice agents, read-aloud101

About a quarter of a cent per 10 seconds of Flash audio, until the price doubles on January 1, 2027.

The two models split the work. Flash TTS is the creative one, for audiobooks, characters and scripted dialogue, in a hundred and thirty languages. Flash Lite is the fast, cheaper one, for dubbing, voice agents and read aloud features at volume. At launch prices, ten seconds of Flash audio costs about a quarter of a cent, and Lite costs less. Google's pricing page says prices for both double on January 1st.

Voices of your own

Gemini 3.8 TTS
010203Voices of your own
02
Voices of your own

Now the voices.

Describe a voice and Gemini makes it

Gemini 3.8 TTS
010203Voices of your own
blog.google · Gemini 3.8 text-to-speech · Sep 23, 2026

Describe a voice in words, and Gemini makes it

Role, accent and character, in over 100 languages and dialects. Or pick from 2,000 ready-made voices.

With Flash TTS, you describe the voice you want in plain words, like a high energy DJ from Melbourne, or a fire breathing dragon, and set its role, accent and character. It works across more than a hundred languages and dialects. If you'd rather not design one, there's a library of more than two thousand ready-made voices, including regional ones like Quebec French and Scots English. Voices you make can be saved and reused, so a character sounds the same next week.

Copy a voice from 30 seconds, with the owner's consent

Gemini 3.8 TTS
010203Voices of your own
blog.google · Gemini 3.8 text-to-speech · Sep 23, 2026

Copy a voice from a 30-second sample, with its owner's spoken consent

Consent checka recording of the owner agreeing must match the voice
Watermarkedevery clip carries a SynthID watermark

You can also copy a real voice, your own or one you have the rights to, from about thirty seconds of audio. Before Gemini will create it, the owner has to record themselves agreeing, and that recording has to match the voice being copied. And every clip these models produce carries Google's SynthID watermark, which you can't hear but which marks it as generated speech. Copied voices also come with C2PA content credentials.

Direct the performance line by line

Gemini 3.8 TTS
010203Voices of your own
blog.google · Gemini 3.8 text-to-speech · Sep 23, 2026

Then direct the performance, line by line

Stage directionsa calm agent, a whispered scene, set per line
Two-speaker scenesfrom one script, with natural turn-taking
Hours of audiowith little drift in the voice

Once you have voices, you direct them. Each line can carry its own direction, from a calm customer service agent to a whispered suspense scene, plus laughs, sighs and little "mm hm" reactions. Two speakers can share one script and take turns naturally. And Google says a voice holds steady across hours of audio, which is what podcasts and audiobooks need.

Who it reaches

Gemini 3.8 TTS
010203Who it reaches
03
Who it reaches

Finally, who can use them.

Where you can use them

Gemini 3.8 TTS
010203Who it reaches
blog.google · Gemini 3.8 text-to-speech · Sep 23, 2026
Developersthe Gemini API and Google AI Studio, now
EveryoneFlash TTS in Gemini Notebook, Flash-Lite in Google Vids
Gemini Enterprisecoming soon

Developers can use both models now, in the Gemini API and in Google AI Studio, which has a new workspace for designing and copying voices. If you don't build anything, you'll hear them in Google's own apps: Flash in Gemini Notebook, and Flash Lite in Google Vids. Gemini Enterprise customers are told it's coming soon.

On the older preview model? Check your scripts

Gemini 3.8 TTS
010203Who it reaches
ai.google.dev · Gemini 3.8 TTS model pages

On the older preview model? Check your scripts before you switch

The new models read your text word for word, so a direction like "Say cheerfully:" may be spoken aloud.

If you already use Gemini's earlier speech model, the 3.1 Flash TTS preview, Google recommends Flash Lite as its replacement, and marks the old one deprecated, with no shutdown date yet. One change catches people out. The new models treat your text strictly as a script, so a direction written inline, like "say cheerfully", may be read out loud. Directions now go in a separate field, and Google's migration guide shows how.

Where to read more

Gemini 3.8 TTS
010203Who it reaches
blog.google · Gemini 3.8 text-to-speech · Sep 23, 2026 · Gemini API release notes · Sep 22, 2026

Speech that's designed, copied with consent, and directed

blog.google › Gemini 3.8 text-to-speech · ai.google.dev › Gemini API release notes

One limit: voice remixing, where you'd take a library voice and adjust it with a prompt, is still marked coming soon. Google's announcement and the Gemini API release notes have the details. That's Gemini 3.8 speech: design a voice, or copy yours.