Caption design colors: hex, rgb(), rgba() or transparent only
Sume caption design colors take hex, rgb(), rgba() or transparent. Other CSS syntax is rejected at request time, so a bad color costs nothing.

What is accepted
The design.colors fields on a Sume caption job accept hex values, rgb() and rgba() functions, or the word transparent. Anything else, such as a named color or another CSS function, is rejected, and per the video captions docs it is not put into the render document. A bad value therefore fails at request time and you do not pay for a wrong render.
Which fields take a color
The six color fields are base, active, stroke, accent, accent_deep and card. A card of null shows no card behind the text. Each field you send overrides one design token of the chosen style, and all other tokens stay as the style defines them. The docs' own example sends only design.colors.active as #22D3EE on black-outline and gets a cyan emphasis with everything else unchanged.
Why it fails early
Typography, placement, phrasing and motion groups work the same way: out-of-range numbers return a 400. The rule behind all of them is the one the docs state: a wrong look should fail when you ask for it, not after a paid render.
Brand palettes often arrive as names or as hsl() values from a design file. Convert them before the request. The script below checks the formats the docs list and prints what it would reject. It is a local guard, not a copy of the server's parser, so the API remains the final check.
import re
HEX = re.compile(r"^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{4}|[0-9a-fA-F]{6}|[0-9a-fA-F]{8})$")
FUNC = re.compile(r"^rgba?\(\s*[\d.%\s,/]+\)$")
def looks_accepted(value):
v = value.strip()
return v == "transparent" or bool(HEX.match(v)) or bool(FUNC.match(v))
palette = {"active": "#22D3EE", "base": "rgba(255,255,255,0.95)",
"stroke": "black", "accent": "hsl(190 90% 50%)", "card": "transparent"}
for name, value in palette.items():
print(name, value, "ok" if looks_accepted(value) else "convert first")What the check prints
Run it and stroke and accent come back as convert first: black is a name, and hsl() is another function. Replace them with #000000 and the hex for that hue.
Three more rules from the same docs
punchandtiktok-greendo not supportdesignat all; they render on a path that reads none of these tokens, so adesignobject there has no effect.- Colors only change the look. The
languagefield remains a speech-to-text hint and never selects a style or a font. - To try a second palette without transcribing again, send
source_caption_idand a newdesign; Sume reuses the word timings it already has.
Cost
A standalone caption job reserves $0.20 for a video of up to 60 seconds. Restyling is still a render, so the price does not change, but the transcription does not run twice. Live prices are on GET /v1/catalog.
Converting a palette in your own code
A design file usually gives colors as names or as hsl(). Convert them once, store the hex values, and send those. A palette is rarely more than six colors, so a one-time lookup table in your repository is enough; there is no need to parse every possible CSS color at request time.
Write the converted value in lowercase or uppercase, whichever your team uses, and keep a comment with the original name so a later designer can see what #1a1a1a was meant to be. Alpha is the one thing hex does not carry in the three-digit form, so use rgba() when a stroke or card needs partial opacity.
After conversion, run the validator from the script above in your build or in a test, so a bad value fails before a caption job is created and not as a 400 in production.
- Keep a lookup table from color name to hex.
- Use
rgba()for partial opacity. - Test the palette in CI, not at request time.
Sources
Related posts
More in Developers
- Captions out of sync with the audio: check STT word times and offsets
Captions running early or late usually trace to an unapplied offset. How Sume STT word times work, which offset to add, and a Python merge that applies it.
- Chapter timestamps for narrated audio from concat segment offsets
Join one TTS file per chapter with timeline audio concat, then turn the returned segments[] start offsets into mm:ss chapter lines with a short Python script.
- Check an Omni edit kept the rest of the clip: video-frames pairs
Compare stills from the source and the edited clip at the same timestamps with video-frames on Sume. A script that submits both extracts, plus what to look for.
- Check an Omni edit's length with video-inspect before a timeline join
An edit should follow the source length. Confirm it with a probe-only video-inspect call before the clip goes into a Timeline render, with a Python read.
Written by Sume