Mixed-language script: one Sume TTS request per language, then concat

A script that switches language mid-way needs one TTS request per language on Sume. Join up to 20 parts with timeline audio at $0.01.

5 min readSume
All posts

Sume's tts_create takes one language and one voice per request, so a script that switches language needs one request per language, with a voice tagged for each. Then join the audio files in order with timeline audio, which concatenates up to 20 parts for a flat $0.01. Do not send a mixed script to a single voice: the language check warns on a mismatch, and the output will not be what you want.

How to split

Break the script at every language switch into parts, and list each as a language, a voice and a transcript. Keep each part a full sentence, since a voice that starts mid-sentence can sound clipped. If you have more than 20 parts, concatenate in two rounds.

Sume limits for a mixed-language read, read 2026-10-03
ItemLimit or price
Transcript per requestUp to 20,000 characters
TTS price$0.0475 per 1,000 characters
Parts per concatUp to 20
Concat price$0.01 flat

Checking the join

Listen at each seam. If the volume or pace differs between voices, use generation_config with speed (0.6 to 1.5) and volume (0.5 to 2) on the quieter part. See timeline audio for the concat shape, and MCP tools and gates for the TTS fields.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume