Switch voiceover language mid-video: one TTS job each, then join

ElevenLabs advertises mid-call language switching for agents. For a produced video on Sume, run one TTS job per language and join them with audio concat.

4 min readSume
All posts

On Sume, give each language its own TTS job with its own language code, then join the clips with the timeline audio concat endpoint. The language field is per request, so a single job should stay in one language. ElevenLabs advertises mid-call language switching for its agents, which is a live-conversation feature rather than something Sume's job API offers.

The result for a produced video is the same: a voiceover that starts in English and continues in Spanish, with no re-synthesis at the seam.

What the vendor feature is

The ElevenLabs v4 Turbo in ElevenAgents post lists mid-call language switching among v4 Turbo's agent features, alongside 90+ languages and gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. That is about a live agent changing language when the caller does. It is not an id on Sume, and Sume's TTS Router catalog is Sonic only.

The Sume recipe

Split the script at each language change. For every segment, send a TTS request with the matching BCP-47 language and the same voice. Keep the voice consistent, and if the voice is not suited to a language, the mismatch warning will ask you to confirm. See the TTS contract for the fields.

Then call timeline audio concat with the clip URLs in order. Concat accepts 1 to 20 parts, joins in the sample domain and does not re-synthesize speech, so the cadence of each clip is kept exactly. Output defaults to wav; choosing mp3 re-adds encoder padding, so stay in wav until the final render.

  • One language per TTS job.
  • Same voice id across jobs.
  • Same output format across parts, so channel layouts match.
  • Concat is a flat per-job charge, listed in the timeline audio docs.

Where it breaks

Do not send a mixed-language paragraph and hope. Names and loan words inside a sentence are fine, but a whole clause in another language needs its own job. If the parts differ in channel layout, concat returns audio_parts_channel_mismatch, which you fix by generating every part with the same output format.

Keep the seams on sentence boundaries. A pause inserted by cutting between sentences sounds natural; a cut mid-sentence sounds like an edit.

Then add captions

Request word timestamps on each TTS job, offset them by the cumulative duration of earlier parts, and feed the combined words into the captions endpoint so the burned-in text follows the audio. The timestamps to captions post walks through the handoff.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume