Mixed-language explainer: two TTS jobs joined into one track
A voice that handles 23 or 90 languages still needs one language per request. Render each language as its own job, then join the takes with timeline audio.

How do you narrate one explainer in two languages with an AI voice? Render each language as its own TTS job with the right language, then join the two takes with timeline audio concat. One request carries one language tag, so do not paste both scripts into a single transcript.
Language coverage is a headline for new voices: Microsoft lists 23 languages for MAI-Voice-2.1 and ElevenLabs lists more than 90 for Eleven v4 (both read 2026-10-04). Coverage tells you a voice can speak a language, not that one request should switch between them.
Set the language on every job
Sume's language field is a BCP-47 tag. If you omit it the request defaults to English, though Hangul-only or kana-only text is detected as Korean or Japanese. A Spanish script sent with no language is read as English, which is the usual cause of a strange accent. If the chosen voice is not suited to the language, the job warns until you set confirm_language_mismatch: true.
Ask for wav on both jobs. Joining mp3 re-adds encoder padding at the seam, while wav joins sample-exact.
Two jobs, one join
Replace the two audio URLs with the artifacts from your finished TTS jobs. Both parts must share a channel layout or the join fails with audio_parts_channel_mismatch.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: explainer-en-es-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_en/take.wav" },
{ "url": "https://media.sume.com/artifacts/artf_es/take.wav" }
]
}'
Use the offsets
The concat result returns segments[] with each part's offset in the joined track. Use them to place captions: shift the second language's word timestamps by its segment offset, then send all words to the caption job.
Check the joined length. TTS audio longer than 1200 seconds fails per job, and timeline audio produces at most 1800 seconds. A short explainer stays well inside both.
Sources
Related posts
More in Use cases
- Naver Clip video: a 9:16 Sume clip with Korean captions
Making a vertical 9:16 clip for Naver Clip? Generate it on Sume, then burn Korean captions with the korean-ad style and language ko.
- Netflix Ads: 10-75 s at 16:9 1920x1080, render 16:9 native
Netflix Ads video specs: 16:9 at 1920x1080, 10 to 75 seconds, H.264. Render 16:9 natively on Sume and sequence clips for longer spots.
- Nonprofit year-end impact recap from five photos for Giving Tuesday
A 50-second year-end recap for a nonprofit from five photos: voice-over, music bed, one Timeline render and captions. $0.4678 on Sume, from docs.
- Digital Omnibus timeline: June 16 to July 27 in order
A dated list of the EU Digital Omnibus steps as reported in 2026, from Parliament on June 16 to entry into force on July 27, with the Art. 50(2) grace date.
Written by Sume