Switch voiceover language mid-video: one TTS job each, then join
ElevenLabs advertises mid-call language switching for agents. For a produced video on Sume, run one TTS job per language and join them with audio concat.

On Sume, give each language its own TTS job with its own language code, then join the clips with the timeline audio concat endpoint. The language field is per request, so a single job should stay in one language. ElevenLabs advertises mid-call language switching for its agents, which is a live-conversation feature rather than something Sume's job API offers.
The result for a produced video is the same: a voiceover that starts in English and continues in Spanish, with no re-synthesis at the seam.
What the vendor feature is
The ElevenLabs v4 Turbo in ElevenAgents post lists mid-call language switching among v4 Turbo's agent features, alongside 90+ languages and gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. That is about a live agent changing language when the caller does. It is not an id on Sume, and Sume's TTS Router catalog is Sonic only.
The Sume recipe
Split the script at each language change. For every segment, send a TTS request with the matching BCP-47 language and the same voice. Keep the voice consistent, and if the voice is not suited to a language, the mismatch warning will ask you to confirm. See the TTS contract for the fields.
Then call timeline audio concat with the clip URLs in order. Concat accepts 1 to 20 parts, joins in the sample domain and does not re-synthesize speech, so the cadence of each clip is kept exactly. Output defaults to wav; choosing mp3 re-adds encoder padding, so stay in wav until the final render.
- One language per TTS job.
- Same voice id across jobs.
- Same output format across parts, so channel layouts match.
- Concat is a flat per-job charge, listed in the timeline audio docs.
Where it breaks
Do not send a mixed-language paragraph and hope. Names and loan words inside a sentence are fine, but a whole clause in another language needs its own job. If the parts differ in channel layout, concat returns audio_parts_channel_mismatch, which you fix by generating every part with the same output format.
Keep the seams on sentence boundaries. A pause inserted by cutting between sentences sounds natural; a cut mid-sentence sounds like an edit.
Then add captions
Request word timestamps on each TTS job, offset them by the cumulative duration of earlier parts, and feed the combined words into the captions endpoint so the burned-in text follows the audio. The timestamps to captions post walks through the handoff.
Sources
Related posts
More in Use cases
- Sync music to avatar scenes: hook, demo, CTA timestamps in the prompt
Match a Lyria bed to a 12-second multi-scene avatar video by writing [0:00-0:03] section markers into the Music Router prompt. A worked example and cost.
- Taboola Realize video ad specs: 15 s motion ad vs 90 s video
Realize (Taboola) lists a 15-second motion ad and a video spec of 6 to 30 s, 90 s max. The sizes, files and how to cut a product clip for each with Sume.
- Tag products in a YouTube Short: the sound rule and a Sume clip
Product tags in a Short: tag in the order products appear, use a Shopping sound or no sound, and block reasons like copyright claims. Plan the clip with Sume.
- Teams translated captions vanish after the meeting: keep them
Microsoft says Teams translated captions and transcripts are only available during the meeting. Caption the recording afterwards with Sume STT and burned cues.
Written by Sume