Edits has 30+ caption styles; Sume has 10 plus design overrides
Instagram says Edits offers 30+ caption styles in multiple languages. Sume Video Captions documents 10 named styles and a design override object. See the table.

Instagram's creators page says Edits caption generation is "available in 30+ styles and multiple languages", with a bulk editor to fix typos and timing (read 2026-10-03). Sume's Video captions documents ten named styles, not thirty: slam, punch, tiktok-green and korean-ad, plus six Hangul identities (black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis). Sume closes part of the gap with a design object that overrides colours, typography, placement, phrasing and motion for one request, but a style count of 30 is not something it offers.
What the ten styles are
Omit style and the wording decides: slam for Latin text, black-outline for Korean. Three of the named styles, slam, punch and tiktok-green, draw in Latin display faces with no Hangul glyphs, so sending Korean to them returns a 400 (caption_hangul_text_latin_style) rather than tofu boxes.
| Style | Use it for |
|---|---|
slam | Default for Latin copy |
punch | Named Latin style; design not supported |
tiktok-green | Named Latin style; design not supported |
korean-ad | Ad karaoke, Korean speech, pair with language: "ko" |
black-outline | Hangul safe default, white fill on thick black outline |
weight-shift | Phrase cards, spoken word takes the weight |
highlight | Accent block sweeps behind the spoken word |
pill-karaoke | Dark pill, colour tracks the voice |
clip-wipe | Words wiped in left to right; clearest at small phone sizes |
editorial-emphasis | Two-line card, final word drops to a larger line |
Recolouring instead of counting styles
A design object merges over the style: colors (base, active, stroke, accent, accent_deep, card), typography, placement, phrasing and motion. Numbers outside their documented range are a 400 at request time, so a bad look fails before it bills. The request below keeps black-outline and changes only the spoken-word colour, which is the cheapest way to get a house look. It uses the speech in your clip, so a silent file fails as caption_no_speech.
import json, os, urllib.request
API = "https://api.sume.com/v1"
def post(path, body, key):
req = urllib.request.Request(
f"{API}{path}",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": key,
},
method="POST",
)
with urllib.request.urlopen(req) as res:
return json.load(res)
job = post(
"/video-captions",
{
"video_url": os.environ["SUME_PUBLIC_VIDEO_URL"],
"style": "black-outline",
"design": {"colors": {"active": "#22D3EE"}},
},
"caption-cyan-001",
)
print(job["request_id"])
Which one to use
If you want a library to browse and tweak with a thumb, Edits is built for that, and it handles per-caption fixes in a bulk editor. If you want a repeatable look applied by a script to many clips, Sume's ten styles plus design tokens are enough, because a brand look is mostly colour, weight and where the line sits.
Pricing is fixed at $0.20 per accepted caption job for videos up to 60 seconds (confirm in GET /v1/catalog). To try another style without a second transcription, pass source_caption_id and Sume reuses the source video and its word timings; a restyle is still a render and still bills.
Building a look with tokens
Because design merges over the chosen style, a consistent brand look is a short JSON object you store once and send with every request. A sensible starting set is colors.active for the spoken word, placement.anchor_ratio to move the line centre as a fraction of frame height, and phrasing.max_words to control how many words appear at once. All of these are optional, and a value outside the documented range is rejected at request time, so you find out before you are billed.
That is a smaller surface than 30 styles, and it favours repeatability over browsing. If a client wants "the same captions as last week", you send the same object. If you want to compare three looks on one clip, transcribe once and restyle with source_caption_id, which reuses the word timings and returns a new render.
A note on languages
Instagram says Edits captions come in multiple languages. Sume's language field is a speech-to-text hint (ko, en and so on); omit it for auto-detect. It never selects the style or the font. Latin copy defaults to slam, Korean to black-outline, and a style you name is rendered as named, even if it is a poor fit for the language, except where the Hangul guard returns a 400.
Sources
Related posts
More in Comparisons
- Edits exports 4K 60fps HDR: Sume Timeline's 2160 px edge and 60 fps
Instagram says Edits exports up to 4K at 60fps with HDR. Sume Timeline 1.0 caps each side at 2160 px and offers 24, 25, 30 or 60 fps. Numbers and a plan call.
- Edits keyframes vs Sume Timeline: stills are static, no zoom
Edits supports keyframes. Sume Timeline holds a still static and ignores a motion field; animate the still first with Video Router image-to-video.
- Eleven v4 90+ languages vs Sonic 3.6 44 languages: which fits yours
ElevenLabs says Eleven v4 covers 90+ languages. Cartesia says Sonic 3.6 covers 44. How to check your language before you pick a TTS model or API.
- Eleven v4 and Sume: a step-by-step feature map for voiceover work
Eleven v4 is not a model id on Sume. Here is what v4 does per ElevenLabs, and which Sume endpoint covers each step of a voiceover pipeline today.
Written by Sume