A UGC-style AI ad in three scenes: hook, demo, call to action
Build a UGC-style spokesperson ad on Sume Avatar 1.0 with ordered video_inputs: a spoken hook, a silent demo beat and a call to action, all in a 9:16 frame.
A UGC-style ad on Sume Avatar 1.0 is three entries in video_inputs: a spoken hook, a silent demo beat and a spoken call to action, rendered at 9:16 with a casual background prompt. Total planned length must stay between 4 and 60 seconds, and all scenes share one avatar and one background.
The structure
The Generate avatar video page (read 2026-10-08) uses this exact pattern as its multi-scene example: a hook of 3 seconds, a 4-second silent demo and a 5-second call to action, for 12 seconds in total.
| Scene id | Voice | Seconds | Job in the ad |
|---|---|---|---|
| hook | text | 3 | A surprised question that stops the scroll |
| demo | silence | 4 | Show the product or result |
| cta | text | 5 | One clear next step |
Writing for the format
UGC ads usually sound like someone talking to a friend, so write short, spoken lines. About three seconds is one short sentence; five seconds is one or two. Keep the hook to a single idea and let the demo beat carry the proof. Use the same background prompt in every scene, for example casual bedroom framing with native UGC lighting, because the current execution expects the backgrounds to resolve to one shared scene.
- Hook: a question or surprise, under 12 words.
- Demo: silence, no narration over the product shot.
- CTA: one action, not two.
- Captions: turn on inline captions, since many viewers watch muted.
Captions and checks
Inline captions burn into the final MP4 using the spoken text. The default style is slam; Hangul styles are required for Korean scripts. A caption-stage failure does not fail the avatar job, so check captions.status in the result. See Video captions for styles.
Run a preview first if the product image matters. It gives one still per scene, and you can regenerate stills cheaply before the full render. The related posts cover cost by tier.
Disclosure
An AI spokesperson that looks like a customer can mislead. Label it as AI where platform rules or local law require, and do not present a generated avatar as a real customer testimonial.
Variations worth testing
Change one element per variant. Keep the avatar and the background the same and test three different hooks, since the hook decides whether a viewer stays. Then take the best hook and test two calls to action. Because each clip is under a minute and uses the same handle, the variants stay comparable.
Preview before you render the set. One approved first frame can save a batch of near-identical rejects, and the preview stills do not depend on the quality tier you choose later. For cost planning, use the per-second tiers described in the pricing posts rather than guessing.
Keep a record of scripts and job ids on your side. Use a different Idempotency-Key for each variant so retries never merge two creative versions.
Sources
Related posts
More in Use cases
- Upscale old footage past 30 seconds: chunk, upscale, reassemble
Sume video upscale takes up to 30 seconds per call at $0.009 per input second. Chunk a 5-minute tape with video-trim, upscale each piece and join with Timeline.
- Used-car listing videos: Omni drafts $0.30, finals $1.50
A 8-second car listing clip on Sume costs $0.30 at 360p and $1.50 at 1080p with Omni Flash. Twenty cars is $30.00 at 1080p. Veo 3.1 prices beside it.
- 20 hooks x 4 cuts = 80 vertical ad clips: cost by Sume model
An 80-clip vertical ad test at 8 seconds and 720p costs $48.00 to $193.60 on Sume. TikTok recommends 9:16 with a minimum of 540x960 for non-Spark ads.
- Walmart's three-flashes-a-second limit and Sume's 24-frame check
Walmart bans strobe and caps flashing at three times a second. Sume has no flash detector; video-frames returns up to 24 named stills to review one second.
Written by Sume