Multilingual video captions: language is a hint, not a font

On Sume video captions, `language` only tells speech-to-text what to expect. Look and font come from `style`, `design` and `font`, never from the language.

4 min readSume
All posts

To caption a Spanish, German, French, Portuguese or Italian video on Sume, send the clip to POST /v1/video-captions and treat language as a transcription hint only. The docs say it plainly: language never selects the style or the font. You choose the look separately with style, design and font.

Descript's changelog announced gradient caption fills and filler-word cleanup in five more languages (Spanish, German, French, Portuguese, Italian). Those are Descript editor features; this page covers what the Sume captions route does with the same languages. Facts below are from the Video captions docs, read 2026-10-01.

What does the language field actually do?

The docs describe it as a speech-to-text hint (ko, en, and so on). Leave it out and the language is detected automatically. It does not restyle anything: a style you name renders as named, whatever language the speech is in.

The only place wording changes the look is an omitted style. Then the wording decides: slam for Latin text, black-outline for Korean. Spanish, German, French, Portuguese and Italian are all written in the Latin alphabet, so an omitted style gives you slam.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-es-001" \
  -d '{
    "video_url": "https://example.com/clip-es.mp4",
    "language": "es",
    "style": "punch"
  }'

Which field controls which part of the caption?

Caption request fields, from the Sume docs read 2026-10-01: https://docs.sume.com/models/video-captions
FieldWhat it controls
languageSpeech-to-text hint; omit for auto-detect
styleThe look and the motion; slam, punch, tiktok-green, korean-ad and the Hangul identities
designPer-request overrides of a style's colours, typography, placement, phrasing and motion
fontA Hangul face for the chosen style; Hangul styles only
script_textScript alignment: the words you want burned in

How do I change the colours or gradient look?

Use design. A style is a set of design tokens and design overrides them for one request, so one key changes one thing. The documented colour fields are base, active, stroke, accent, accent_deep and card, given as hex, rgb()/rgba() or transparent. The docs do not describe a gradient fill, so do not plan around one. Note that design is not supported on punch or tiktok-green; see caption design overrides.

What does a multilingual batch cost?

The caption job price is on the rate card and in GET /v1/catalog; the docs state $0.20 for videos up to 60 seconds under the current fixed estimate. Confirm it there before a large batch.

Should I pass language for every video?

If you know it, pass it; if the batch mixes languages, omit it and let detection run. Either way, pick style once for the whole batch so every language shares one look. If the caption text drifts from what was said, pass script_text for alignment (see Video captions).

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume