TikTok-style hook, demo, CTA in one 9:16 avatar video call
Build a 12-second 9:16 hook, demo and CTA clip with Sume Avatar video_inputs, a silent demo beat and tiktok-green captions in a single request.
You can build a hook, demo and call-to-action clip in one request with Sume Avatar 1.0 by sending ordered video_inputs: a spoken hook, a silent demo beat, and a spoken CTA. The default aspect_ratio is already 9:16, and inline captions with style: "tiktok-green" burn the words in after generation.
Below is a 12-second plan with the exact fields, and what the docs say about limits. Nothing here is a claim about how TikTok will rank it; it is a clean way to produce a short vertical clip.
What shape does the request take?
Avatar video takes a ready avatar by avatar_handle plus exactly one of script or video_inputs. Total planned duration must land between 4 and 60 seconds inclusive. Each scene has an id, a voice and a background. A spoken scene uses voice.type: "text" with exactly one of script or input_text; a silent beat uses voice.type: "silence" with a required duration and no text.
Execution currently supports one resolved avatar per final video, and expects scene backgrounds to resolve to one shared scene, so keep the same background prompt across all three.
| Scene id | Voice | Seconds | Job |
|---|---|---|---|
| hook | text | 3 | One line that names the payoff |
| demo | silence | 4 | A beat with no speech while the product shows |
| cta | text | 5 | One sentence telling the viewer the next step |
What does the request look like?
This uses the documented fields only. Replace avatar_handle with one of your own ready avatars:
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hook-demo-cta-001" \
-d '{
"avatar_handle": "YOUR_AVATAR_HANDLE",
"aspect_ratio": "9:16",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "This one change fixed my morning.", "duration": 3},
"background": {"type": "prompt", "prompt": "Bright kitchen, handheld framing"}},
{"id": "demo", "voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Bright kitchen, handheld framing"}},
{"id": "cta", "voice": {"type": "text", "input_text": "Try it for a week and see.", "duration": 5},
"background": {"type": "prompt", "prompt": "Bright kitchen, handheld framing"}}
],
"captions": {"enabled": true, "style": "tiktok-green", "language": "auto"}
}'What do the captions do in a silent scene?
Inline captions use the spoken script and video_inputs text, so the silent demo scene has no words to burn. That is by design: the caption follows speech. If you want text over the demo beat, use the standalone video captions route afterwards with cues carrying text, start and end in seconds, which skips speech-to-text.
Know the style limits too. tiktok-green does not accept design overrides, since it renders on a path that reads none of those tokens; use slam if you need to tweak colours. A Korean script on a Latin style such as tiktok-green returns 400 caption_hangul_text_latin_style, so Korean speech needs a Hangul style.
What happens if captions fail?
Caption stage failures soft-fail: the avatar job can still succeed with a clean video_url and captions.status=failed. Check that field before you publish, and re-caption the clean video with the standalone route if needed. Poll the job through jobs and results rather than assuming the first response is final.
Resolution is currently 720p, and the scenes are generated, so review a first-frame preview before a full render if the look matters. See first-frame previews.
Checklist
Before you submit:
- Scene durations add up to between 4 and 60 seconds.
- Silent scenes carry
durationand no script. - All three backgrounds describe the same scene.
- Captions style matches the script language.
- You check
captions.statuson the result.
Sources
Related posts
More in Use cases
- TikTok in-feed ad minimum 516 kbps bitrate: what Sume gives you
TikTok's non-Spark in-feed ad spec asks for 516 kbps or more, 500 MB or less. Sume's trim and filter docs list no bitrate control, so here is the check.
- TikTok in-feed ad file types: MP4, MOV, MPEG, 3GP or AVI?
TikTok's non-Spark in-feed ads accept five file types, Spark only .mp4 or .mov. Sume trim returns MP4, so a trimmed AI clip fits both.
- TikTok Auto-Selection for lead gen: what to upload as a UGC clip
TikTok added Auto-Selection to Lead Generation campaigns. It draws on brand uploads, creator content and AI content, so upload several 9:16 clips, not one.
- TikTok Multilingual tool: translated ad audio and captions
TikTok's Multilingual tool translates an ad's audio and adds captions per device language, but only in Smart+ campaigns. What it covers and what Sume adds.
Written by Sume