Why does the weight-shift caption style look flat with another font?
weight-shift and korean-ad animate the wght axis, which only Pretendard has. With any other font they keep colour and scale emphasis but lose the weight travel.

The weight shift disappears because it needs a variable font, and in Sume's caption renderer only Pretendard carries the wght axis. If you set font to another face, weight-shift and korean-ad still burn, with the colour and scale emphasis intact, but the spoken word no longer gets heavier.
This is documented rather than a bug. The Video captions page says these two styles animate the weight axis, and that on a static face they keep their colour and scale emphasis and lose the weight travel.
Which combinations keep the effect
Omitting font keeps each style's own face, so the default combination is the one that animates. The faces are all open-licence fonts shipped with the renderer. An unknown font name is rejected, never replaced, so a caption cannot fall back to a face you did not choose.
A Latin style cannot take a Hangul font at all. Naming one returns 400 with caption_font_requires_hangul_style, and sending Korean copy to slam, punch or tiktok-green returns caption_hangul_text_latin_style.
| Style | Default face | Animates the wght axis | With another font set |
|---|---|---|---|
| korean-ad | Pretendard | Yes | Colour and scale emphasis only |
| weight-shift | Pretendard | Yes | Colour and scale emphasis only |
| highlight | Pretendard | Not documented | Face swapped |
| pill-karaoke | Pretendard | Not documented | Face swapped |
| editorial-emphasis | Pretendard lead line, Black Han Sans emphasis line | Not documented | Lead line swapped, emphasis line stays Black Han Sans |
| black-outline | Do Hyeon | Not documented | Face swapped |
| clip-wipe | Do Hyeon | Not documented | Face swapped |
How to keep both a custom look and the weight shift
If you only wanted a different colour or size, do not change the font. Use the design field, which overrides colours, typography, placement, phrasing and motion for one request. A design.colors.active of #22D3EE, for example, changes the active-word colour and leaves the face alone. Note that design is not supported on punch or tiktok-green.
If you do want a particular face, pick the style whose effect does not depend on weight, such as clip-wipe with jua, rather than expecting weight-shift to animate it.
Testing a font without paying for transcription twice
Once you have one captioned job, re-burn it under a new style or font by sending source_caption_id instead of video_url. Sume reuses the source video and its word timings, so no second speech-to-text runs, though a restyle is still billed as a render. That makes a three-way comparison of the same clip cheap to set up; the restyle post walks through it.
Sources
Related posts
More in Media tools
- Where to break subtitle lines: BBC rules and Sume caption cues
The BBC says one sentence per subtitle and no article-noun splits. Sume sentence segments cover the first rule; cues with a line break cover the rest.
- Which Sume media tools take a 3-minute Reel and a 5-minute TikTok
A 3-minute Reel is 180 s and a 5-minute TikTok is 300 s. Which Sume media tools accept that source length, which stop short, and the one fix for each limit.
- Who is speaking in subtitles: BBC colours, dashes, and Sume cues
WCAG 1.2.2 wants speaker identification in captions. The BBC prefers colour, then dashes or labels. What Sume's burned-in captions can do and a dash script.
- YouTube A/B tests skip Shorts: make two caption-style variants
YouTube's A/B tool does not cover Shorts. Re-burn one Short in a second caption style with source_caption_id, then publish both and compare by hand.
Written by Sume