Holiday ad captions on Avatar 1.0: slam, punch or tiktok-green?
Avatar 1.0 burns captions inline in slam (default), punch or tiktok-green. No separate billed caption job, and a caption failure keeps the clean video.
Turn on inline captions in the avatar video request and pick slam (the default), punch or tiktok-green for Latin-script holiday ads. The captions are burned into the final MP4 after generation, they do not create a separate billed caption job, and if the caption stage fails, the avatar job can still succeed with a clean video.
Most holiday clips are watched with the sound off. Captions are not decoration, they are the ad.
The settings
The avatar video docs define the captions object with enabled, style and language. The text comes from the spoken script or the video_inputs text. Sume never adds captions to preview stills, so you will not see them in a preview. Captions take the same settings as standalone video captions: style, an optional font, a language hint and script_text.
| Style | Use it for |
|---|---|
| slam (default) | Latin-script speech; the style you get when you do not choose |
| punch | Latin-script speech |
| tiktok-green | Latin-script speech |
| korean-ad | Hangul karaoke for Korean speech |
| weight-shift, black-outline, highlight, pill-karaoke, clip-wipe, editorial-emphasis | Hangul identities |
Two behaviors to build around
The docs warn about one failure. If a Korean script uses slam, punch or tiktok-green, Sume rejects the request with 400 caption_hangul_text_latin_style, since those font faces render Hangul as tofu. Sume does not change the style for you. If you localize a campaign, pick the style by script, not by taste.
The other point is the soft failure. If the caption stage fails, the job can still succeed with a clean primary video_url and captions.status=failed. Always read captions.status on the result. If it failed, run the clean video through the standalone video captions API, which takes a public HTTPS video_url.
The captions object
The request body is the same one you use for a script-based clip, with a captions object added.
{
"avatar_handle": "product_host",
"script": "Free gift wrap on every order until Sunday.",
"aspect_ratio": "9:16",
"quality": "standard",
"captions": {
"enabled": true,
"style": "punch",
"language": "auto"
}
}Write for the muted viewer
For a holiday run, keep the script lines short, since captions show phrase by phrase and a long line reads poorly on a phone. Put the price or the deadline in the spoken line so it is also in the caption, and then check the burned text against your price list. A caption is generated from speech or the script, so a misheard number would appear on screen.
If you must change the price later, re-caption the clean clip rather than re-rendering the avatar. The standalone captions API can take authored cues with text, start and end times, and no speech-to-text, which is useful when the text must be exact.
Placement, phrasing and consistency
Think about where the caption sits. Platform interfaces cover the bottom of a vertical video with buttons, the username and the description, so a caption that sits low can be hidden. Check one clip on a real phone in the app where the ad will run before you queue a batch. If the caption collides with the interface, the fix is in the script and the framing, not in a retry loop.
Keep a holiday script to one idea per line. A line such as the offer, the deadline and the code can be three short sentences, and each reads as its own caption phrase. A single long sentence with the offer buried at the end is read poorly with or without captions.
Style choice is a brand decision. Pick one of the three Latin styles for the whole campaign so the ads look like a series, and test a second only as a deliberate A/B. Changing the caption style between ads in the same flight makes the test noisy, since you then compare two things at once.
What captions cost
Billing is simple. Inline captions do not add a separate billed job on the avatar request, so a clip costs the avatar rate for its seconds: $0.184 per second on standard, $0.245 on plus and $0.55 on max for a no-product clip. A 12-second plus clip is $2.94 with or without captions. The standalone captions route is a different job with its own price, which you pay only if you re-caption a clean clip later.
Sources
Related posts
More in Sume Avatar 1.0
- Map avatar video scene ids to onboarding steps with scene_previews
Give each video_inputs scene a stable id and Sume returns it in scene_previews with start, end and duration, so one clip can drive per-step chapters.
- One voice across 23 languages: MAI-Voice-2.1 vs a Sume avatar voice
MAI-Voice-2.1 keeps one voice across 23 languages. A Sume voice has one primary language and a 409 guard on mismatch. What that means for avatars.
- Pick a stock avatar by avoid_for and brand_safety_notes, not looks
Sume's avatar catalog returns profile metadata with best_for, avoid_for, brand_safety_notes and casting_notes. Read them before you cast a presenter for a clip.
- Griffin-Lite 26 of 54 Turing result: what the sample size says
Tavus reports 26 of 54 callers fooled by Griffin-Lite after a one-minute call. A 95% interval is about 35% to 61%. Python computes it, and says what to claim.
Written by Sume