Caption 40 clips in six languages with no language hint: $8
Leave `language` off and Sume's caption job detects it. Forty clips of up to 60 seconds cost $8.00 at $0.20 each; here is the loop and the style trap.

Yes: omit language and Sume's caption job lets speech-to-text find the language itself, so a folder of clips in six languages can go through one loop at $0.20 a clip, $8.00 for 40 clips of up to 60 seconds. The only trap is the style, which Sume picks from the caption text when you do not send one.
What the docs say about language
Sume's video captions page defines language as a speech-to-text hint (ko, en and so on) and says that if you do not send it, speech-to-text finds the language automatically. It also says language never selects the style or the font; it only tells speech-to-text which language to expect.
That matters for a mixed batch. Microsoft's MAI-Transcribe-2 page highlights automatic language detection across 60 languages for its streaming model (read 2026-10-05), which is aimed at live audio. For recorded clips you do not need a streaming model. A hint is optional and a batch with no hints is valid.
The style trap in a mixed batch
If you do not send style, the text of the captions sets it: slam for Latin text and black-outline for Korean. That is the safe path for a batch where you do not know each clip's language in advance. If you hard-code style: "slam" for the whole batch, a Korean clip returns 400 caption_hangul_text_latin_style, because slam, punch and tiktok-green use Latin display faces with no Hangul glyphs.
The page documents Latin and Hangul handling. It does not list fonts for other scripts, so run one clip per script you expect, and look at the burn before you queue 40.
| Clip speech | Style Sume picks | Language hint sent | Price |
|---|---|---|---|
| English, Spanish, German, French, Portuguese or Italian | slam (Latin text) | None; auto-detected | $0.20, up to 60 s |
| Korean | black-outline | None; auto-detected | $0.20, up to 60 s |
Any language, but style: slam forced | slam | None | 400 on Hangul text |
| A silent clip | Not applicable | None | Fails with caption_no_speech |
The loop
Each clip needs a public HTTPS video_url that Sume can fetch, and a stable Idempotency-Key so a retry never creates a second paid job.
import os, requests
API = "https://api.sume.com/v1/video-captions"
KEY = os.environ["SUME_API_KEY"]
urls = [u.strip() for u in open("clips.txt") if u.strip()]
submitted = {}
for i, url in enumerate(urls):
r = requests.post(
API,
headers={
"Authorization": f"Bearer {KEY}",
"Idempotency-Key": f"cap-mixed-{i:03d}",
},
json={"video_url": url},
timeout=60,
)
r.raise_for_status()
submitted[url] = r.json()
print(len(submitted), "jobs, about $", round(len(submitted) * 0.20, 2))What the $8 does and does not cover
The $0.20 price applies to videos of up to 60 seconds under the current fixed estimate, and the page says to confirm the live price in GET /v1/catalog. A 90-second clip is outside that statement, so split it or check the catalog before you count on 20 cents.
Failures are not billed as a surprise: a clip with no speech fails with caption_no_speech and a next_action of use_overlay_captions. For those, send cues with text, start and end instead of speech-to-text.
Fixing words after detection
Auto-detect can mishear a brand name. Send script_text to align your own wording onto the speech-to-text timings, or restyle a finished caption with source_caption_id and corrected words, which reuses the stored word timings and does not run speech-to-text a second time. A restyle is still a render, so the price does not change.
A sampling plan before the full run
Do not send all 40 at once on the first day. Pull one clip per language, which is six jobs and $1.20, and check three things in the finished video: that the detected words are right, that the style fits the script, and that the line breaks look sensible at phone width. Only then queue the remaining 34.
Keep the job id next to each source URL in your own table. If a later review finds one clip with a misheard product name, you can restyle that single caption from source_caption_id with corrected words instead of paying for a second transcription of the whole batch.
Sources
Related posts
More in Developers
- Caption cue limits: 400 characters, 200 cues, 60 seconds
A caption cue takes 1-400 characters, start of 0 or more, end above start and 60 s or less; a request holds 1-200 cues. Validate locally before the $0.20 job.
- Caption job 400: words, cues, segments, script_text are exclusive
The video captions API accepts only one wording source: words, cues or segments (which skip STT), or script_text (aligned onto STT). Sending two returns a 400.
- Chain a 30-second render, trim and captions: three jobs, three keys
Generate, trim and caption an AI clip through the Sume API as three separate jobs, each with its own Idempotency-Key, status poll and result read.
- Check a callback_url is public HTTPS before you submit a Sume video
Sume's callback_url must be public HTTPS. A Python pre-check rejects http, embedded credentials, unresolvable hosts and private addresses before you pay.
Written by Sume