Gemini 3.8 TTS caps multi-speaker at two voices: a three-voice scene
Gemini 3.8 Flash TTS supports up to two speakers in one multi-speaker request. For a three-voice scene on Sume, make one TTS take per speaker and join them.

The Gemini speech generation guide says multi-speaker output supports up to 2 speakers with prebuilt voices. For a scene with three characters, you need another route. On Sume, send one TTS take per line with the right voice, then join the takes in order with a Timeline audio concat. Cost is per character, so the number of voices does not change the price.
What the vendor page says
The Gemini API speech guide lists the current models gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Both accept text-only input, and the multi-speaker mode is limited to two speakers using prebuilt voices. A third character means a second request and your own stitching.
The Sume approach
Sume TTS has no speaker limit because each request is a single voice. You split the script by speaker, make one take per line, and concatenate.
- Split the script into lines tagged with a speaker.
- Send each line to
POST /v1/tts-1.0/generatewith that speaker'svoice.id. Ask for wav. - Send all the takes, in script order, as
parts[]toPOST /v1/timeline-1.0/audiowithoperation: concat. A job takes 1 to 20 parts. - Use the returned
segments[]offsets to place any visuals.
Keeping the voices distinct
Give each character its own voice id, and keep generation_config per character, so the same person sounds the same on every line. Use a frozen preset for each speaker so speed, volume and language do not drift. For a pause between speakers, add a short silence part rather than padding the text.
Stay on the same sample rate and format for every take. Mixed rates make the join harder to reason about; wav at 44100 Hz keeps everything sample-exact.
Cost and limits
TTS is $0.0475 per 1,000 characters, so a 600-character scene is the same whether one voice or three read it. Timeline audio concat is $0.01 flat per job. Each concat output must stay within 1,800 seconds and 20 parts, so a long dialogue is joined in chapters. See the concat limits.
Many short lines add up in requests, not dollars. Group consecutive lines from the same speaker into one take to stay under 20 parts.
Check the join
After the concat, play the joined file once end to end. Listen for level jumps between voices; gain_db on a part (range -60 to 12) lets you balance a quiet voice. Check that the segments[] offsets match your expected line lengths before you place visuals on them.
If one line needs redoing, regenerate only that take and rerun the concat, rather than the whole scene.
Sources
Related posts
More in Use cases
- Translated text in an image runs longer: check fit before you publish
When you translate the text in an image into a longer language, lines may not fit. A fit checklist for Ideogram 4.5 edits on Sume, and when to redraw in code.
- Gift guide by budget: three short Timeline renders beat one long one
Split a holiday gift guide into under-$25, $50 and $100 brackets, one short Timeline render each, and plan all three free first. Each is $0.10 a started minute.
- GitHub README banner with light and dark versions from Sume
Generate two 3:1 banners with a 1536x512 custom size, commit them to the repo, and switch them with a picture element and prefers-color-scheme in the README.
- Godot: import a Sume image as a texture, WebP or PNG
Godot imports PNG, JPEG and WebP. Drop a Sume image into res://, pick Lossless or VRAM Compressed, and use 1024x1024 where power-of-two sizes help.
Written by Sume