English captions that highlight the spoken word: which style does it

Which Sume Latin caption style colors the word being spoken, where design.colors.active and active_scale apply, and a Python request that burns the result.

4 min readSume
All posts

For English speech, punch and tiktok-green color the word being spoken inside a short phrase (red-pink #ff3f57 and green #39E508), but they refuse the design object, so those colors are fixed. slam accepts design but shows one word at a time, and design.colors.active tints its keyword picks rather than every spoken word. A job is $0.20 for a video up to 60 seconds.

Read on 2026-10-03: the Sume video-captions docs page, plus the caption renderer and API validation code in the Sume repository that the page describes.

Which English style animates the active word?

The docs list three Latin styles: slam, punch and tiktok-green. They behave differently, and design (colors, typography, phrasing, motion overrides) is only accepted by one of them.

Latin caption styles and the active word, read 2026-10-03
StyleWhat the viewer seesdesign overrides
slamOne word on screen at a time. Gold #FFD700 marks keyword picks: not the first word, not a stopword, five letters or fewer.Accepted
punchA phrase page that switches about every 1.2 s; the spoken word turns #ff3f57.Rejected with a 400
tiktok-greenThe same paging; the spoken word turns #39E508.Rejected with a 400

What do design.colors.active and active_scale change?

On slam, colors.active recolors the keyword picks. Its defaults for active_weight and active_scale equal the base look (400 and 1), and the slam template I read applies the active color and motion timings, not a scale or weight jump. If you want a weight and scale change on the spoken word, those tokens (typography.active_weight, typography.active_scale, range 0.5 to 2) matter most on phrase-card styles such as korean-ad and weight-shift, which default to weight 800 to 900 and scale 1.06.

Those phrase-card styles are the Hangul identities (weight-shift, highlight, pill-karaoke and others). The docs steer Korean speech to them. The API only refuses Hangul wording on a Latin style, not Latin wording on a Hangul style, but I did not verify how well their faces draw English, so render one test clip first.

Runnable request: slam with a cyan keyword color

Set SUME_API_KEY and use a public HTTPS clip. The script submits, polls the job and prints the result JSON.

import os, time, requests

API = "https://api.sume.com"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
    "video_url": "https://example.com/clip.mp4",
    "style": "slam",
    "design": {"colors": {"active": "#22D3EE"}},
}
r = requests.post(f"{API}/v1/video-captions", json=body,
                  headers={**AUTH, "Idempotency-Key": "active-word-demo-001"})
r.raise_for_status()
job = r.json()["data"]
while True:
    s = requests.get(job["status_url"], headers=AUTH).json()["data"]
    if s["terminal"]:
        break
    time.sleep(s.get("next_poll_after_seconds") or 3)
print(requests.get(job["result_url"], headers=AUTH).json())

How do I try another style without paying for speech-to-text again?

Pass source_caption_id with a new style instead of video_url. Sume reuses the source video and the word timings it already has, so no second transcription runs, but a restyle is still a billed render. Swap style to tiktok-green and drop design for the fixed green highlight.

Phrasing knobs (max_words, max_chars, pause_seconds) are inert on slam because it is one word at a time, so use them on a phrase-card style.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume