Coffee roaster holiday video: caption in your brand colour with design
Caption a roaster's holiday video in a brand colour with `design.colors.active`, then restyle twice without a second transcription: three looks for $0.60.

To caption a coffee roaster's holiday subscription video in the brand's colour, send the clip to POST /v1/video-captions with a style and a design.colors.active hex value, then restyle it with source_caption_id so speech-to-text does not run again. A standalone caption job is $0.20 for a clip up to 60 seconds, and a restyle is still a render, so three looks cost $0.60.
Small roasters usually have their own footage: the roaster talking over the drum, a pour, a gift box being closed. The captions are the part that has to look like the brand, and the API lets you move colour, weight and placement without re-editing the video.
Three looks, three renders
Each accepted caption job reserves and captures $0.20 for clips of at most 60 seconds, under the current fixed estimate. Inline Avatar Video captions are a separate add-on and create no video_caption resource. Confirm the live price in GET /v1/catalog.
| Job | Request shape | Sume price |
|---|---|---|
| First look: style slam, brand gold on the spoken word | video_url + style + design | $0.20 |
| Second look: black-outline, cream accent | source_caption_id + style + design | $0.20 |
| Third look: placement and phrasing tuned | source_caption_id + design | $0.20 |
| Total | $0.60 |
Which styles take design
design overrides tokens of a style for one request, and each field is optional. The groups are colors (base, active, stroke, accent, accent_deep, card), typography, placement, phrasing and motion. Colours are hex, rgb(), rgba() or transparent. Any other CSS syntax is rejected, and a number outside its range is a 400, so a bad look fails before you pay for a render.
Two styles are the exception: punch and tiktok-green do not support design, because they render on a path that reads none of those tokens. Choose slam or one of the Hangul identities when you need your own colours. For Latin text with no style, the default is slam.
style picks the look and the motion, design changes that look, font picks the face (Hangul styles only) and language only hints speech-to-text. Language never selects the style or the font, so an English roaster video with a Spanish voice-over still needs you to choose a style on purpose.
A sensible house workflow is to fix one brand design object in your repo, send it with every caption job, and change only style per campaign. Because a restyle costs the same $0.20 as a first render, the cheapest test is a single clip, three looks, one pick, and then the rest of the catalogue.
Brand gold on the spoken word
import os
import uuid
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": f"roaster-caption-{uuid.uuid4()}"}
body = {
"video_url": "https://media.sume.com/artifacts/example/roaster.mp4",
"style": "slam",
"language": "en",
"design": {"colors": {"active": "#C8963E"}},
}
r = requests.post("https://api.sume.com/v1/video-captions",
headers=H, json=body, timeout=60)
print(r.status_code, r.json().get("next_action"))
Restyle without transcribing twice
Then restyle with { "source_caption_id": "...", "style": "black-outline" }. Sume reuses the first caption's source video and its word timings, so the second and third looks skip transcription. Send words along with it only to correct the text, and send only one of script_text, words, cues and segments.
If the owner's spoken product names are rare words, send script_text so the burned text matches your spelling: Sume keeps the speech-to-text timings as the clock and aligns your script to them, and a failed alignment returns script_alignment_mismatch or script_alignment_failed with the next action simplify_script_text_or_omit.
Limits to know
video_url must be a public HTTPS clip that Sume can fetch, and localhost, private-network, signed or provider task URLs are rejected. A silent clip fails with caption_no_speech; for a b-roll-only cut, send cues with text, start and end instead. The finished result gives you the status, the style and the captioned video_url as an artifact; raw transcripts and signed source URLs are not part of the public contract, so keep your own copy of any wording you want to reuse. SRT uploads are not accepted either, and phrase-level text goes through cues or segments. Korean text on slam is refused with caption_hangul_text_latin_style, which the Hangul font post explains. For price-line variants with cues instead of speech, see twelve Black Friday price variants.
Sources
Related posts
More in Media tools
- Color-grade AI video on Sume: lut3d is off, tone filters stay
Sume video filter has a dim op, a crop op and an allowlisted ffmpeg filtergraph. lut3d is not on it. What a grade can and cannot do, per the docs.
- Comedy skit music: one bed and two stings with AI music
Score a 45-second comedy skit with one light bed and two short stings: three Music Router generations at $0.125 each plus a Timeline render, $0.475 in total.
- Continue a Wan 3.0 clip: trim the tail, send it as a reference video
Wan 3.0 takes reference videos on Sume (up to 5, 15 s total). To continue a scene, cut the last 3 seconds with video trim and use them as the next reference.
- Corporate onboarding video music: a 5-minute ducked bed
Add a 5-minute instrumental bed under onboarding narration with duck_db in Timeline: one generation and a 5-minute render, $0.625 on Sume.
Written by Sume