Localize one video ad into many languages with Sume batches

Two ways to localize one ad on Sume: a bulk Format batch with one item per language, or burn translated cue captions onto the existing video for $0.20 per job.

6 min readSume
All posts

To localize one ad into many languages on Sume, choose by what has to change. If only the on-screen words change, burn translated captions onto the finished video with the video-captions API. If the spoken script, the scene or the voice must change too, run a Format once per language in a bulk batch with one item for each language.

Path A: translated captions on the same video

POST /v1/video-captions accepts authored overlay text as cues or segments, each with text, start and end in seconds. With those forms Sume does not run speech-to-text and burns your text at your times. You send only one of script_text, words, cues and segments.

Pick a style that has glyphs for the target script. A Latin style given Hangul text is rejected with caption_hangul_text_latin_style instead of rendering missing-glyph boxes, so choose a matching style for the languages you ship. language is only a speech-to-text hint and never selects the style or the font.

The source language of the speech matters only for transcription. If you already have a translated script, send it as cues with your own timings and speech-to-text never runs. If you have only the original speech, run the first job on the original video, and then restyle the same caption job for look changes, not language changes, since source_caption_id reuses the original words.

Path B: one Format run per language

When the script itself changes, queue one item per language on a Format and put the language in input or in the instruction. Each item is a full agent turn that produces its own video, so it costs a full run and takes minutes, but nothing about the output is limited to captions.

Bind the same output_schema to every item so the results line up. A nullable locale_note field is a safe place for the run to say what it could not do, and video as a SumeMediaFile gives you one URL per language.

Which path to choose

Localization paths on Sume (read 2026-10-07)
QuestionCaptions (Path A)Bulk Format (Path B)
What changesBurned-in text onlyAnything the run can regenerate
UnitOne caption jobOne Format run per item, 1 to 100 per queue
Price signal$0.20 per job for a video of 60 seconds or less (current fixed estimate)Capped by generation_spend_cap_usd on each item
TimeA render jobA full agent turn per language
Restyle latersource_caption_id reuses the transcript; no second speech-to-textNew run, or previous_run_id to continue one

Cost of the captions path

At $0.20 per accepted job, ten languages on a clip of 60 seconds or less come to $2.00 in caption jobs. The live rate is on GET /v1/catalog, so read it before you plan a large run. Path B has no per-run price in the docs, only the cap you set, so size it from a trial run.

Checks before you publish

  • Have a native reader approve the cues. Sume burns the text you send, it does not check meaning.
  • Keep cue text short. The captions page says its clearest style at small phone sizes is clip-wipe.
  • For Path B, read filled_by and output_error on each receipt before counting a language as delivered.
  • Give each caption job its own Idempotency-Key, one per language, so a retry does not pay twice.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume