Caption 40 clips in six languages with no language hint: $8

Leave `language` off and Sume's caption job detects it. Forty clips of up to 60 seconds cost $8.00 at $0.20 each; here is the loop and the style trap.

5 min readSume
All posts

Yes: omit language and Sume's caption job lets speech-to-text find the language itself, so a folder of clips in six languages can go through one loop at $0.20 a clip, $8.00 for 40 clips of up to 60 seconds. The only trap is the style, which Sume picks from the caption text when you do not send one.

What the docs say about language

Sume's video captions page defines language as a speech-to-text hint (ko, en and so on) and says that if you do not send it, speech-to-text finds the language automatically. It also says language never selects the style or the font; it only tells speech-to-text which language to expect.

That matters for a mixed batch. Microsoft's MAI-Transcribe-2 page highlights automatic language detection across 60 languages for its streaming model (read 2026-10-05), which is aimed at live audio. For recorded clips you do not need a streaming model. A hint is optional and a batch with no hints is valid.

The style trap in a mixed batch

If you do not send style, the text of the captions sets it: slam for Latin text and black-outline for Korean. That is the safe path for a batch where you do not know each clip's language in advance. If you hard-code style: "slam" for the whole batch, a Korean clip returns 400 caption_hangul_text_latin_style, because slam, punch and tiktok-green use Latin display faces with no Hangul glyphs.

The page documents Latin and Hangul handling. It does not list fonts for other scripts, so run one clip per script you expect, and look at the burn before you queue 40.

What happens with no style and no language (Sume video captions page, read 2026-10-05)
Clip speechStyle Sume picksLanguage hint sentPrice
English, Spanish, German, French, Portuguese or Italianslam (Latin text)None; auto-detected$0.20, up to 60 s
Koreanblack-outlineNone; auto-detected$0.20, up to 60 s
Any language, but style: slam forcedslamNone400 on Hangul text
A silent clipNot applicableNoneFails with caption_no_speech

The loop

Each clip needs a public HTTPS video_url that Sume can fetch, and a stable Idempotency-Key so a retry never creates a second paid job.

import os, requests

API = "https://api.sume.com/v1/video-captions"
KEY = os.environ["SUME_API_KEY"]
urls = [u.strip() for u in open("clips.txt") if u.strip()]

submitted = {}
for i, url in enumerate(urls):
    r = requests.post(
        API,
        headers={
            "Authorization": f"Bearer {KEY}",
            "Idempotency-Key": f"cap-mixed-{i:03d}",
        },
        json={"video_url": url},
        timeout=60,
    )
    r.raise_for_status()
    submitted[url] = r.json()

print(len(submitted), "jobs, about $", round(len(submitted) * 0.20, 2))

What the $8 does and does not cover

The $0.20 price applies to videos of up to 60 seconds under the current fixed estimate, and the page says to confirm the live price in GET /v1/catalog. A 90-second clip is outside that statement, so split it or check the catalog before you count on 20 cents.

Failures are not billed as a surprise: a clip with no speech fails with caption_no_speech and a next_action of use_overlay_captions. For those, send cues with text, start and end instead of speech-to-text.

Fixing words after detection

Auto-detect can mishear a brand name. Send script_text to align your own wording onto the speech-to-text timings, or restyle a finished caption with source_caption_id and corrected words, which reuses the stored word timings and does not run speech-to-text a second time. A restyle is still a render, so the price does not change.

A sampling plan before the full run

Do not send all 40 at once on the first day. Pull one clip per language, which is six jobs and $1.20, and check three things in the finished video: that the detected words are right, that the style fits the script, and that the line breaks look sensible at phone width. Only then queue the remaining 34.

Keep the job id next to each source URL in your own table. If a later review finds one clip with a misheard product name, you can restyle that single caption from source_caption_id with corrected words instead of paying for a second transcription of the whole batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume