150-language subtitles: which scripts Sume captions document
A translation model can output 150 languages, but Sume documents Latin and Hangul caption styles. Test other scripts on a short clip before a batch.

A model that translates into 150 languages does not mean every one of them renders on screen. Sume's caption docs describe Latin styles and Hangul styles. For anything else, such as Thai, Arabic or Hindi, the docs promise nothing, so burn a three-cue test clip and look at it before you translate a whole library.
Index-Translate's README lists 150 languages for translation. That is a claim about text output. Whether a caption renderer has a font that covers a script is a separate question, and the answer for Sume is only documented for two script families.
What do the Sume docs actually say?
The caption page names slam, punch and tiktok-green as Latin styles, and a set of Hangul styles including black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis, plus korean-ad. The font field picks a Hangul face and applies to Hangul styles only. Korean text on a Latin style is rejected with a 400 and the code caption_hangul_text_latin_style.
The docs do not list Thai, Arabic, Devanagari or other scripts, and they do not say they are unsupported either. That is the reason to test.
| Script | Documented in Sume caption docs | What to do |
|---|---|---|
| Latin (English, Spanish, French and so on) | Yes, Latin styles | Use slam, punch or tiktok-green |
| Hangul (Korean) | Yes, Hangul styles and fonts | Use a Hangul style; a Latin style returns a 400 |
| Thai, Arabic, Devanagari, others | Not documented | Burn a 3-cue test and inspect every glyph |
How do you run the test?
Right-to-left scripts and scripts with joined or stacked letters are the usual trouble spots in any renderer, so judge those by eye and not by a status code. A successful job only tells you that a file was produced.
- Pick three cues that include the hardest text: the longest line, one with digits and punctuation, and one with a brand name.
- Post them as
cueson a short public HTTPS clip, so the test costs one standalone caption job at most. - Open the result and check for missing glyphs, boxes, broken joins between letters, and wrong reading direction.
- Only after it looks right, translate and burn the full set.
What is the fallback if a script does not render?
Keep the translated text and ship it a different way: as a sidecar subtitle file your player renders, or on a Latin transliteration line burned in. The bilingual subtitle walkthrough shows a two-line approach that fits when one line is Latin.
For the languages Sume documents, the recipe in the cue-translation post applies unchanged. A standalone caption job is priced at $0.20 for videos up to 60 seconds under the current estimate.
What should the test clip contain?
Use a clip that is cheap to burn, with a plain background that will show missing glyphs clearly. Three cues are enough, but pick them with care: one with the longest word you expect, one with numerals, and one with a line that wraps. If the script is written right to left, make sure at least one cue contains mixed text, such as a Latin brand name inside the line, because ordering bugs usually appear at the boundary between scripts.
Save the test output next to your notes with the date and the style name. When the docs change, you can re-run the same three cues and see whether anything improved.
Sources
Related posts
More in Developers
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
- Trigger.dev public tokens: keep the Sume key server-side
Trigger.dev v4.6.2 hardened authorization for public tokens. Whatever token your browser holds, a Sume API key must never be one of them. Here is the split.
- TTS 400: Provide exactly one of transcript_source or transcript
This 400 means the TTS body had both transcript and transcript_source, or neither. Send exactly one, plus a voice, and re-run.
- TTS 400 asking for voice.id or avatar_id / avatar_handle: the fix
The TTS call needs a voice: set voice.id, or a top-level avatar_id or avatar_handle. Without one the request is rejected before any audio is made.
Written by Sume