Mixed-language explainer: two TTS jobs joined into one track

A voice that handles 23 or 90 languages still needs one language per request. Render each language as its own job, then join the takes with timeline audio.

6 min readSume
All posts

How do you narrate one explainer in two languages with an AI voice? Render each language as its own TTS job with the right language, then join the two takes with timeline audio concat. One request carries one language tag, so do not paste both scripts into a single transcript.

Language coverage is a headline for new voices: Microsoft lists 23 languages for MAI-Voice-2.1 and ElevenLabs lists more than 90 for Eleven v4 (both read 2026-10-04). Coverage tells you a voice can speak a language, not that one request should switch between them.

Set the language on every job

Sume's language field is a BCP-47 tag. If you omit it the request defaults to English, though Hangul-only or kana-only text is detected as Korean or Japanese. A Spanish script sent with no language is read as English, which is the usual cause of a strange accent. If the chosen voice is not suited to the language, the job warns until you set confirm_language_mismatch: true.

Ask for wav on both jobs. Joining mp3 re-adds encoder padding at the seam, while wav joins sample-exact.

Two jobs, one join

Replace the two audio URLs with the artifacts from your finished TTS jobs. Both parts must share a channel layout or the join fails with audio_parts_channel_mismatch.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: explainer-en-es-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_en/take.wav" },
      { "url": "https://media.sume.com/artifacts/artf_es/take.wav" }
    ]
  }'

Use the offsets

The concat result returns segments[] with each part's offset in the joined track. Use them to place captions: shift the second language's word timestamps by its segment offset, then send all words to the caption job.

Check the joined length. TTS audio longer than 1200 seconds fails per job, and timeline audio produces at most 1800 seconds. A short explainer stays well inside both.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume