Which caption style wins? Test slam, punch, tiktok-green

Burn three caption styles onto one clip for $0.60 and compare them. Latin text works on all three; Korean text is rejected on these styles. Runnable loop.

3 min readSume
All posts

Run the same clip through three caption jobs, one per style, and post whichever holds attention: it costs $0.60 at $0.20 per standalone caption job. slam, punch and tiktok-green are three of the documented styles, and each job returns a captioned video from the same source.

The loop

Caption once from the video, then restyle by source_caption_id, so speech to text runs a single time. This sends the requests and prints each id; request_id is the field the job envelope returns.

import json, os, urllib.request

def post(body, key):
    req = urllib.request.Request(
        "https://api.sume.com/v1/video-captions",
        data=json.dumps(body).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json", "Idempotency-Key": key})
    return json.load(urllib.request.urlopen(req))

first = post({"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
              "style": "slam", "language": "en"}, "ab-cap-slam")
print("slam", first.get("request_id"))
# wait here until the first caption job is finished
for style in ["punch", "tiktok-green"]:
    r = post({"source_caption_id": first["request_id"], "style": style}, "ab-cap-" + style)
    print(style, r.get("request_id"))

Know the limits

  • Latin display faces have no Hangul glyphs. Korean text sent to slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style; use black-outline or korean-ad for Korean speech.
  • punch and tiktok-green do not support design overrides, so tune their look by choosing a different style instead.
  • A clip with no speech fails with caption_no_speech; send cues instead.

Cost of testing

Standalone caption jobs for videos up to 60 s, read 2026-10-08
Styles testedJobsCost
11$0.20
22$0.40
33$0.60
55$1.00

Why restyle by id

The docs say a restyle reuses the word timings the source caption already has, so speech to text does not run a second time, and the price does not change because a restyle is still a render. A style test therefore stays predictable at $0.20 per variant. The restyle needs the first caption job to have finished, so poll it before the loop continues.

Reading the test

Judge the three versions on a phone at arm's length with the sound off. A good caption style is legible against the busiest part of your footage, moves just enough to guide the eye, and does not cover faces. Choose one style per channel and keep it, because a consistent look reads as a brand. If none of the three suits, the design override on slam lets you change colors and placement for a single request without touching the other values.

Related posts

More in Use cases

All Use cases posts

Written by Sume