Localized caption batch: eight QA checks before you ship
Eight checks for a batch of burned caption videos in several languages: script coverage, language hint, timing, key per job, and cost. Sume job behavior.

Before you ship a batch of captioned videos in several languages, check eight things: the script is covered by a documented face, the language hint matches the speech, the source has audible speech, the cue times fit the translated text, each job has its own Idempotency-Key, no job reuses another's timings, the cost matches the number of languages, and every output is watched once. Each check below maps to behavior on the Video captions page.
The guidance is limited to what the docs state; the translations are yours.
1. Does every language have a documented face?
The docs list Latin styles and Hangul styles. Korean copy on a Latin style returns 400 (caption_hangul_text_latin_style). For other scripts nothing is documented, so test one clip. See Japanese, Chinese or Arabic captions.
2. Is the language hint right?
language is a speech-to-text hint such as ko or en; omit it for automatic detection. It never selects the style or font. Use it when the speech is clear and you know the language, and omit it when a clip mixes languages.
3 to 5. Speech, timing, and keys
A silent clip fails as caption_no_speech with next_action: use_overlay_captions; send cues with text, start, and end instead. Translated text often needs more time on screen, so check each cue against its window. Use a different Idempotency-Key per language and per deliberate rerun, and reuse a key only for the same payload.
| Check | Docs behavior | Error or field |
|---|---|---|
| Hangul on Latin style | Rejected before billing | caption_hangul_text_latin_style |
| Silent clip | Fails, suggests overlay captions | caption_no_speech |
| Script does not align | Typed failure; omit script_text | script_alignment_mismatch |
| Out-of-range design number | 400 at request time | design fields |
| Dubbed file | Needs its own job | video_url, not source_caption_id |
6 to 8. Reuse, cost, and the final watch
source_caption_id reuses the source video and word timings, so it is for restyling the same clip, not a dub. See caption a dubbed video. Each accepted job is $0.20 for video up to 60 seconds, so N languages are N times $0.20.
Last, watch every output once. The docs say the caption stage returns a video, and a render that looks wrong still bills as a good one, so a skim of each language is the cheapest check you have.
Sources
Related posts
More in Media tools
- Loop background music under a long video with Timeline 1.0
A 30-second bed under a 90-second video: use soundtrack.loop, fade_out_seconds and duck_db in a Sume Timeline 1.0 render, and plan first for free.
- LTX-2.5 Alpha Gen: video mattes without a green screen
LTX-2.5 Alpha Gen turns RGB clips of up to 145 frames into grayscale mattes. What it needs, where it stops, and what Sume's RMBG route does instead.
- LTX-2.5 Restore LoRA for archive footage vs Sume video upscale
LTX-2.5 Restore cleans and colorizes damaged archive clips with tiled 8-step runs. Sume Video Upscale enlarges by 1.1 to 4 times and does not colorize.
- LTX-2.5 SDR-to-HDR LoRA: ACEScg EXR and HLG, not an MP4 API
LTX-2.5 SDR-To-HDR outputs scene-linear ACEScg EXR frames plus an HLG master from an SDR clip. Sume's video routes return MP4. Where each fits.
Written by Sume