TikTok-style hook, demo, CTA in one 9:16 avatar video call

Build a 12-second 9:16 hook, demo and CTA clip with Sume Avatar video_inputs, a silent demo beat and tiktok-green captions in a single request.

5 min readSume
All posts

You can build a hook, demo and call-to-action clip in one request with Sume Avatar 1.0 by sending ordered video_inputs: a spoken hook, a silent demo beat, and a spoken CTA. The default aspect_ratio is already 9:16, and inline captions with style: "tiktok-green" burn the words in after generation.

Below is a 12-second plan with the exact fields, and what the docs say about limits. Nothing here is a claim about how TikTok will rank it; it is a clean way to produce a short vertical clip.

What shape does the request take?

Avatar video takes a ready avatar by avatar_handle plus exactly one of script or video_inputs. Total planned duration must land between 4 and 60 seconds inclusive. Each scene has an id, a voice and a background. A spoken scene uses voice.type: "text" with exactly one of script or input_text; a silent beat uses voice.type: "silence" with a required duration and no text.

Execution currently supports one resolved avatar per final video, and expects scene backgrounds to resolve to one shared scene, so keep the same background prompt across all three.

Hook, demo, CTA plan (read 2026-10-02)
Scene idVoiceSecondsJob
hooktext3One line that names the payoff
demosilence4A beat with no speech while the product shows
ctatext5One sentence telling the viewer the next step

What does the request look like?

This uses the documented fields only. Replace avatar_handle with one of your own ready avatars:

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hook-demo-cta-001" \
  -d '{
    "avatar_handle": "YOUR_AVATAR_HANDLE",
    "aspect_ratio": "9:16",
    "video_inputs": [
      {"id": "hook", "voice": {"type": "text", "script": "This one change fixed my morning.", "duration": 3},
       "background": {"type": "prompt", "prompt": "Bright kitchen, handheld framing"}},
      {"id": "demo", "voice": {"type": "silence", "duration": 4},
       "background": {"type": "prompt", "prompt": "Bright kitchen, handheld framing"}},
      {"id": "cta", "voice": {"type": "text", "input_text": "Try it for a week and see.", "duration": 5},
       "background": {"type": "prompt", "prompt": "Bright kitchen, handheld framing"}}
    ],
    "captions": {"enabled": true, "style": "tiktok-green", "language": "auto"}
  }'

What do the captions do in a silent scene?

Inline captions use the spoken script and video_inputs text, so the silent demo scene has no words to burn. That is by design: the caption follows speech. If you want text over the demo beat, use the standalone video captions route afterwards with cues carrying text, start and end in seconds, which skips speech-to-text.

Know the style limits too. tiktok-green does not accept design overrides, since it renders on a path that reads none of those tokens; use slam if you need to tweak colours. A Korean script on a Latin style such as tiktok-green returns 400 caption_hangul_text_latin_style, so Korean speech needs a Hangul style.

What happens if captions fail?

Caption stage failures soft-fail: the avatar job can still succeed with a clean video_url and captions.status=failed. Check that field before you publish, and re-caption the clean video with the standalone route if needed. Poll the job through jobs and results rather than assuming the first response is final.

Resolution is currently 720p, and the scenes are generated, so review a first-frame preview before a full render if the look matters. See first-frame previews.

Checklist

Before you submit:

  • Scene durations add up to between 4 and 60 seconds.
  • Silent scenes carry duration and no script.
  • All three backgrounds describe the same scene.
  • Captions style matches the script language.
  • You check captions.status on the result.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume