Edits has 30+ caption styles; Sume has 10 plus design overrides

Instagram says Edits offers 30+ caption styles in multiple languages. Sume Video Captions documents 10 named styles and a design override object. See the table.

4 min readSume
All posts

Instagram's creators page says Edits caption generation is "available in 30+ styles and multiple languages", with a bulk editor to fix typos and timing (read 2026-10-03). Sume's Video captions documents ten named styles, not thirty: slam, punch, tiktok-green and korean-ad, plus six Hangul identities (black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis). Sume closes part of the gap with a design object that overrides colours, typography, placement, phrasing and motion for one request, but a style count of 30 is not something it offers.

What the ten styles are

Omit style and the wording decides: slam for Latin text, black-outline for Korean. Three of the named styles, slam, punch and tiktok-green, draw in Latin display faces with no Hangul glyphs, so sending Korean to them returns a 400 (caption_hangul_text_latin_style) rather than tofu boxes.

Sume Video Captions styles, read 2026-10-03
StyleUse it for
slamDefault for Latin copy
punchNamed Latin style; design not supported
tiktok-greenNamed Latin style; design not supported
korean-adAd karaoke, Korean speech, pair with language: "ko"
black-outlineHangul safe default, white fill on thick black outline
weight-shiftPhrase cards, spoken word takes the weight
highlightAccent block sweeps behind the spoken word
pill-karaokeDark pill, colour tracks the voice
clip-wipeWords wiped in left to right; clearest at small phone sizes
editorial-emphasisTwo-line card, final word drops to a larger line

Recolouring instead of counting styles

A design object merges over the style: colors (base, active, stroke, accent, accent_deep, card), typography, placement, phrasing and motion. Numbers outside their documented range are a 400 at request time, so a bad look fails before it bills. The request below keeps black-outline and changes only the spoken-word colour, which is the cheapest way to get a house look. It uses the speech in your clip, so a silent file fails as caption_no_speech.

import json, os, urllib.request

API = "https://api.sume.com/v1"

def post(path, body, key):
    req = urllib.request.Request(
        f"{API}{path}",
        data=json.dumps(body).encode(),
        headers={
            "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
            "Content-Type": "application/json",
            "Idempotency-Key": key,
        },
        method="POST",
    )
    with urllib.request.urlopen(req) as res:
        return json.load(res)

job = post(
    "/video-captions",
    {
        "video_url": os.environ["SUME_PUBLIC_VIDEO_URL"],
        "style": "black-outline",
        "design": {"colors": {"active": "#22D3EE"}},
    },
    "caption-cyan-001",
)
print(job["request_id"])

Which one to use

If you want a library to browse and tweak with a thumb, Edits is built for that, and it handles per-caption fixes in a bulk editor. If you want a repeatable look applied by a script to many clips, Sume's ten styles plus design tokens are enough, because a brand look is mostly colour, weight and where the line sits.

Pricing is fixed at $0.20 per accepted caption job for videos up to 60 seconds (confirm in GET /v1/catalog). To try another style without a second transcription, pass source_caption_id and Sume reuses the source video and its word timings; a restyle is still a render and still bills.

Building a look with tokens

Because design merges over the chosen style, a consistent brand look is a short JSON object you store once and send with every request. A sensible starting set is colors.active for the spoken word, placement.anchor_ratio to move the line centre as a fraction of frame height, and phrasing.max_words to control how many words appear at once. All of these are optional, and a value outside the documented range is rejected at request time, so you find out before you are billed.

That is a smaller surface than 30 styles, and it favours repeatability over browsing. If a client wants "the same captions as last week", you send the same object. If you want to compare three looks on one clip, transcribe once and restyle with source_caption_id, which reuses the word timings and returns a new render.

A note on languages

Instagram says Edits captions come in multiple languages. Sume's language field is a speech-to-text hint (ko, en and so on); omit it for auto-detect. It never selects the style or the font. Latin copy defaults to slam, Korean to black-outline, and a style you name is rendered as named, even if it is a poor fit for the language, except where the Hangul guard returns a 400.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume