Which caption style wins? Test slam, punch, tiktok-green
Burn three caption styles onto one clip for $0.60 and compare them. Latin text works on all three; Korean text is rejected on these styles. Runnable loop.

Run the same clip through three caption jobs, one per style, and post whichever holds attention: it costs $0.60 at $0.20 per standalone caption job. slam, punch and tiktok-green are three of the documented styles, and each job returns a captioned video from the same source.
The loop
Caption once from the video, then restyle by source_caption_id, so speech to text runs a single time. This sends the requests and prints each id; request_id is the field the job envelope returns.
import json, os, urllib.request
def post(body, key):
req = urllib.request.Request(
"https://api.sume.com/v1/video-captions",
data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json", "Idempotency-Key": key})
return json.load(urllib.request.urlopen(req))
first = post({"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"style": "slam", "language": "en"}, "ab-cap-slam")
print("slam", first.get("request_id"))
# wait here until the first caption job is finished
for style in ["punch", "tiktok-green"]:
r = post({"source_caption_id": first["request_id"], "style": style}, "ab-cap-" + style)
print(style, r.get("request_id"))Know the limits
- Latin display faces have no Hangul glyphs. Korean text sent to
slam,punchortiktok-greenreturns400 caption_hangul_text_latin_style; useblack-outlineorkorean-adfor Korean speech. punchandtiktok-greendo not supportdesignoverrides, so tune their look by choosing a different style instead.- A clip with no speech fails with
caption_no_speech; sendcuesinstead.
Cost of testing
| Styles tested | Jobs | Cost |
|---|---|---|
| 1 | 1 | $0.20 |
| 2 | 2 | $0.40 |
| 3 | 3 | $0.60 |
| 5 | 5 | $1.00 |
Why restyle by id
The docs say a restyle reuses the word timings the source caption already has, so speech to text does not run a second time, and the price does not change because a restyle is still a render. A style test therefore stays predictable at $0.20 per variant. The restyle needs the first caption job to have finished, so poll it before the loop continues.
Reading the test
Judge the three versions on a phone at arm's length with the sound off. A good caption style is legible against the busiest part of your footage, moves just enough to guide the eye, and does not cover faces. Choose one style per channel and keep it, because a consistent look reads as a brand. If none of the three suits, the design override on slam lets you change colors and placement for a single request without touching the other values.
Related posts
More in Use cases
- Vertical 9:16 Seedance ad at 720p: $1.512 on Mini, $4.6224 on 2.5
An 8-second 720p vertical ad costs from the Mini price to the Seedance 2.5 price on Sume. All four tiers at 6, 8, 12 and 15 s, and the 720x1280 frame size.
- Wine tasting invite video: Wan 3.0, 9:16, 12 s, audio reference
A 12-second vertical invite on Wan 3.0 costs $0.75 at 480p and $1.50 at 720p on Sume, and takes up to 5 audio references. Age gates and limits.
- Yoga class promo video: three beats in one 10-second clip
A studio promo in 9:16 with Gemini Omni Flash: three timecoded beats in 10 seconds, calm native sound in the prompt, $1.25 at 720p on Sume.
- YouTube AI disclosure: a clip of someone giving advice they never gave
YouTube's disclosure page lists making someone appear to give advice they never gave. Its examples in a table, and what Sume trims and captions change.
Written by Sume