Localize a 30-second product video into 3 languages: swap the spine
Keep one edit and swap only the voice: three Timeline renders ($0.10) plus three caption jobs ($0.20) cost $0.90 on Sume for English, Spanish and French.

To localize a 30-second product video into three languages, keep the picture and swap only the Timeline spine: render once per language with a different audio.url, then burn captions per language with POST /v1/video-captions. Each render is $0.10 and each caption job is $0.20, so English, Spanish and French cost $0.90 in total, with the voice-over files made elsewhere.
The mistake in most localization pipelines is re-editing the video for each language. The cut does not depend on the language, only the length of the lines does, so the picture can stay as it is and the audio can change.
The same approach scales to a long list: ten languages is ten renders at $0.10 and ten caption jobs at $0.20, or $3.00, still on one edit. The work moves from the video editor to the voice-over, which is where the language-specific effort belongs.
Same slots, new spine
The Timeline is built around this split. audio.url is one spine of up to 1800 seconds, audio.duration_seconds sets the output length, and video[] slots are placed on that spine by start. If all three voice-overs are 30 seconds, the same slot list works for all three renders. If one language runs 2 seconds longer, change duration_seconds and extend the last slot; coverage may stop at most 0.5 seconds before the end of the spine.
Keep the voice lines at the same length by writing to a word budget. A 30-second slot holds roughly the same number of words in each language only by accident, and Spanish and French usually need more words than English, so write the shorter line first.
Check that each voice file is a media.sume.com artifact or asset before you render. Import voice files with POST /v1/media-imports first, because the Timeline refuses off-host URLs at admit with unsupported_media_source.
| Language | Render | Captions | Sume price |
|---|---|---|---|
| English | audio.url = en.wav | language en, slam | $0.30 |
| Spanish | audio.url = es.wav | language es, slam | $0.30 |
| French | audio.url = fr.wav | language fr, slam | $0.30 |
| Total | $0.90 |
Captions, per language
Captions go on after the render, because they are burned onto a finished clip. Send each rendered MP4 to the caption endpoint with language set as a hint to speech-to-text (es, fr); language does not select the style or the font. Latin-script text defaults to slam, which has Latin faces, so Spanish and French are covered by the default.
Do not send Hangul text to slam, punch or tiktok-green: the API rejects it with caption_hangul_text_latin_style. That is not a concern for these three, but it matters when you add a fourth language.
If speech-to-text mis-hears a product name in one language, send script_text for that clip, or send words with corrected word times.
Plan the lengths before you record. Ask the voice talent for a take under 30 seconds in each language, trim to the spine, and run the free POST /v1/timeline-1.0/plan once per language: it returns duration_seconds, segment_count and billable_minutes without creating a job, so a 61-second spine that would bill two minutes shows up before it costs anything.
Three renders from one slot list
import os
import uuid
import requests
API = "https://api.sume.com/v1/timeline-1.0/render"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
slots = [("CLIP_0", 0, 8), ("CLIP_1", 8, 8), ("CLIP_2", 16, 8), ("CLIP_3", 24, 6)]
video = [{"source_url": os.environ[k], "start": s, "duration": d}
for k, s, d in slots]
for lang in ("EN", "ES", "FR"):
body = {"audio": {"url": os.environ[f"VOICE_{lang}"], "duration_seconds": 30},
"video": video}
r = requests.post(API, json=body, timeout=60,
headers={**H, "Idempotency-Key": f"loc-{lang}-{uuid.uuid4()}"})
print(lang, r.status_code, r.json().get("request_id"))
Run it and what it does not do
Poll the three jobs, then send each video_url to the caption endpoint. Mind the idempotency key: a stable key per language and per release, such as loc-ES-2026-11, makes a retry safe, whereas a random one can create a second billable job.
What this does not do: it does not translate anything, and it does not change what is on screen. Any on-screen text in the clips stays in the language it was shot in, so build text plates as stills and swap those with compose, at $0.02 a shot, or leave text out of the footage entirely. A localised plate for each language is a good use of cue-based price text. Review each version with a fluent speaker before it goes live.
Sources
Related posts
More in Use cases
- Localize app store screenshot captions: 6 locales x 5 shots, $1.13
Translate the caption on 5 marketing screenshots into 6 locales with Ideogram 4.5 edits on Sume: 30 calls, $1.125 at low, and why real UI gets re-captured.
- Localize one 30-second ad into six languages: Wan 3.0 cost on Sume
Six language versions of a 30 second ad on Wan 3.0 cost $22.50 at 720p or $45.00 at 1080p on Sume. Draft all six at 480p for $11.25 first.
- Localize one ad into 14 languages: TTS, captions and render for $4.62
One 450-character ad in 14 shared languages costs $0.42 TTS, $2.80 captions and $1.40 render on Sume: $4.62 total. The loop uses one job per language.
- Localize a YouTube thumbnail into 3 languages for $0.11 on Sume
One finished thumbnail, three Ideogram 4.5 edits through POST /v1/images: Spanish, Portuguese, German at low quality for about $0.11, with the prompt and code.
Written by Sume