Coffee roaster holiday video: caption in your brand colour with design

Caption a roaster's holiday video in a brand colour with `design.colors.active`, then restyle twice without a second transcription: three looks for $0.60.

5 min readSume
All posts

To caption a coffee roaster's holiday subscription video in the brand's colour, send the clip to POST /v1/video-captions with a style and a design.colors.active hex value, then restyle it with source_caption_id so speech-to-text does not run again. A standalone caption job is $0.20 for a clip up to 60 seconds, and a restyle is still a render, so three looks cost $0.60.

Small roasters usually have their own footage: the roaster talking over the drum, a pour, a gift box being closed. The captions are the part that has to look like the brand, and the API lets you move colour, weight and placement without re-editing the video.

Three looks, three renders

Each accepted caption job reserves and captures $0.20 for clips of at most 60 seconds, under the current fixed estimate. Inline Avatar Video captions are a separate add-on and create no video_caption resource. Confirm the live price in GET /v1/catalog.

Brand look test on one 40-second clip (read 2026-10-05)
JobRequest shapeSume price
First look: style slam, brand gold on the spoken wordvideo_url + style + design$0.20
Second look: black-outline, cream accentsource_caption_id + style + design$0.20
Third look: placement and phrasing tunedsource_caption_id + design$0.20
Total$0.60

Which styles take design

design overrides tokens of a style for one request, and each field is optional. The groups are colors (base, active, stroke, accent, accent_deep, card), typography, placement, phrasing and motion. Colours are hex, rgb(), rgba() or transparent. Any other CSS syntax is rejected, and a number outside its range is a 400, so a bad look fails before you pay for a render.

Two styles are the exception: punch and tiktok-green do not support design, because they render on a path that reads none of those tokens. Choose slam or one of the Hangul identities when you need your own colours. For Latin text with no style, the default is slam.

style picks the look and the motion, design changes that look, font picks the face (Hangul styles only) and language only hints speech-to-text. Language never selects the style or the font, so an English roaster video with a Spanish voice-over still needs you to choose a style on purpose.

A sensible house workflow is to fix one brand design object in your repo, send it with every caption job, and change only style per campaign. Because a restyle costs the same $0.20 as a first render, the cheapest test is a single clip, three looks, one pick, and then the rest of the catalogue.

Brand gold on the spoken word

import os
import uuid
import requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
     "Idempotency-Key": f"roaster-caption-{uuid.uuid4()}"}
body = {
    "video_url": "https://media.sume.com/artifacts/example/roaster.mp4",
    "style": "slam",
    "language": "en",
    "design": {"colors": {"active": "#C8963E"}},
}
r = requests.post("https://api.sume.com/v1/video-captions",
                  headers=H, json=body, timeout=60)
print(r.status_code, r.json().get("next_action"))

Restyle without transcribing twice

Then restyle with { "source_caption_id": "...", "style": "black-outline" }. Sume reuses the first caption's source video and its word timings, so the second and third looks skip transcription. Send words along with it only to correct the text, and send only one of script_text, words, cues and segments.

If the owner's spoken product names are rare words, send script_text so the burned text matches your spelling: Sume keeps the speech-to-text timings as the clock and aligns your script to them, and a failed alignment returns script_alignment_mismatch or script_alignment_failed with the next action simplify_script_text_or_omit.

Limits to know

video_url must be a public HTTPS clip that Sume can fetch, and localhost, private-network, signed or provider task URLs are rejected. A silent clip fails with caption_no_speech; for a b-roll-only cut, send cues with text, start and end instead. The finished result gives you the status, the style and the captioned video_url as an artifact; raw transcripts and signed source URLs are not part of the public contract, so keep your own copy of any wording you want to reuse. SRT uploads are not accepted either, and phrase-level text goes through cues or segments. Korean text on slam is refused with caption_hangul_text_latin_style, which the Hangul font post explains. For price-line variants with cues instead of speech, see twelve Black Friday price variants.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume