Gemini 3.8 TTS caps multi-speaker at two voices: a three-voice scene

Gemini 3.8 Flash TTS supports up to two speakers in one multi-speaker request. For a three-voice scene on Sume, make one TTS take per speaker and join them.

5 min readSume
All posts

The Gemini speech generation guide says multi-speaker output supports up to 2 speakers with prebuilt voices. For a scene with three characters, you need another route. On Sume, send one TTS take per line with the right voice, then join the takes in order with a Timeline audio concat. Cost is per character, so the number of voices does not change the price.

What the vendor page says

The Gemini API speech guide lists the current models gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Both accept text-only input, and the multi-speaker mode is limited to two speakers using prebuilt voices. A third character means a second request and your own stitching.

The Sume approach

Sume TTS has no speaker limit because each request is a single voice. You split the script by speaker, make one take per line, and concatenate.

  • Split the script into lines tagged with a speaker.
  • Send each line to POST /v1/tts-1.0/generate with that speaker's voice.id. Ask for wav.
  • Send all the takes, in script order, as parts[] to POST /v1/timeline-1.0/audio with operation: concat. A job takes 1 to 20 parts.
  • Use the returned segments[] offsets to place any visuals.

Keeping the voices distinct

Give each character its own voice id, and keep generation_config per character, so the same person sounds the same on every line. Use a frozen preset for each speaker so speed, volume and language do not drift. For a pause between speakers, add a short silence part rather than padding the text.

Stay on the same sample rate and format for every take. Mixed rates make the join harder to reason about; wav at 44100 Hz keeps everything sample-exact.

Cost and limits

TTS is $0.0475 per 1,000 characters, so a 600-character scene is the same whether one voice or three read it. Timeline audio concat is $0.01 flat per job. Each concat output must stay within 1,800 seconds and 20 parts, so a long dialogue is joined in chapters. See the concat limits.

Many short lines add up in requests, not dollars. Group consecutive lines from the same speaker into one take to stay under 20 parts.

Check the join

After the concat, play the joined file once end to end. Listen for level jumps between voices; gain_db on a part (range -60 to 12) lets you balance a quiet voice. Check that the segments[] offsets match your expected line lengths before you place visuals on them.

If one line needs redoing, regenerate only that take and rerun the concat, rather than the whole scene.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume