Four-language product video from recorded voiceovers: $1.20

Reuse one edit for English, Korean, Spanish and French. Swap the audio spine per language: four Sume renders and four caption jobs cost $1.20.

5 min readSume
All posts

To ship one product video in English, Korean, Spanish and French, keep the video slots and swap the audio spine per language on Timeline 1.0, then burn captions in each language. Four renders at $0.10 plus four caption jobs at $0.20 cost $1.20 for a 40-second video (read 2026-10-08). This assumes you already have the four voiceover files, recorded or produced elsewhere, and import them to Sume first.

What changes per language

Only the audio block changes. The video[] slots, transitions and output stay the same. The catch is length: a Spanish read is rarely the same duration as the English one, and Timeline requires that coverage not stop more than 0.5 seconds before the end of the spine. Set audio.duration_seconds to each file's real length and extend the last slot's duration to match, or the render refuses the program.

Per-language changes as of 2026-10-08
FieldEnglishKoreanSpanish / French
audio.urlen.wavko.waves.wav / fr.wav
audio.duration_secondsits own lengthits own lengthits own length
Last video slot durationadjustedadjustedadjusted
Caption styleslam (default for Latin)korean-ad or black-outlineslam, language hint es / fr

The cost

Timeline 1.0 charges $0.10 per ceil(output minute), so a 40-second video bills one minute in every language. Captions are $0.20 per job for clips of up to 60 seconds. This table prices Sume's steps only; it does not include producing the voiceovers.

Four languages, 40-second video as of 2026-10-08
LineArithmeticCost
Timeline renders4 x $0.10$0.40
Caption jobs4 x $0.20$0.80
Total$0.40 + $0.80$1.20

Two caption traps

Latin caption styles (slam, punch, tiktok-green) have no Hangul glyphs. Sending Korean text to one returns 400 caption_hangul_text_latin_style, and Sume does not switch styles for you. Use korean-ad or another Hangul identity for the Korean file and leave font unset unless you want a different Hangul face.

The language field is only a hint for speech-to-text. It never picks the style or the font. If you have the translated script, pass script_text so Sume aligns your wording to the speech timings; a mismatch raises script_alignment_mismatch and you can drop the field to burn the transcribed text instead.

When one voiceover has two parts

If each language has an intro line and a body file, use audio.parts[]: up to 20 gapless slices joined in the sample domain with no re-synthesis. That avoids a separate join job. A standalone join via timeline audio is a flat $0.01 per job if you want to keep the merged file.

{
  "audio": {
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/es-intro.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/es-body.wav" }
    ],
    "duration_seconds": 41
  },
  "video": [
    { "source_url": "https://media.sume.com/artifacts/artf_demo/edit.mp4", "start": 0, "duration": 41 }
  ]
}

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume