UGC video prompt for AI avatars: what goes in each field
A UGC video prompt for an AI avatar is three inputs, not one: the creator's words, a scene prompt for the setting and light, and a product image.
A UGC video prompt for an AI avatar isn't one block of text. A talking UGC clip has three parts that go in different places: the words the creator says, the setting and light, and the product. In Sume's Avatar 1.0 API, the words go in script (or each scene's voice), the setting in a scene prompt or photo, and the product in product_image. The person's look was fixed when you created the avatar.
Field rules come from Sume's Generate avatar video docs and the Sume API reference, read on 2026-09-28; anything called current behavior is read from Sume's code. For the lines themselves, see UGC ad script for AI avatar videos; for the person, how to write an AI avatar prompt.
What goes in each part of a UGC avatar prompt?
Map each thing you would type into a one-box generator to its field:
| Part of the clip | Field | What to write |
|---|---|---|
| The person | avatar_handle | Nothing: the look comes from the avatar you created |
| The words | script, or a scene's voice with type: "text" | Only the lines the creator says |
| A pause for a demo | A scene's voice with type: "silence" and a duration | No words; the seconds it lasts |
| Setting and light | scene (a prompt or a photo), or each scene's background | Where it's filmed and how it's lit |
| The product | product_image | A public HTTPS image URL; leave it out for no product |
| The frame | aspect_ratio | 9:16 by default; also 1:1, 3:4, 4:3, 16:9 |
How do I write the scene prompt?
Name the place, the light, and the framing in one short phrase. The docs' own UGC example uses "Casual bedroom framing, native UGC lighting". Other phrases in that style:
- Kitchen counter, morning window light, close framing.
- Parked car interior, daylight through the windshield.
- Bathroom mirror, warm vanity light, chest-up framing.
- Use one setting per video: in current code, scenes with different backgrounds are refused with a
400, so repeat the same prompt on each scene. A real room works too, as a public HTTPS photo inscene.
What does a full UGC avatar request look like?
A spoken hook, a silent demo beat, and a spoken call to action, each with the same background, plus the product:
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ugc-serum-v1" \
-d '{
"avatar_handle": "ugc_creator",
"product_image": "https://example.com/serum.png",
"video_inputs": [
{ "id": "hook",
"voice": { "type": "text", "script": "Okay, this is the serum everyone keeps asking me about." },
"background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
{ "id": "demo",
"voice": { "type": "silence", "duration": 4 },
"background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
{ "id": "cta",
"voice": { "type": "text", "script": "The link is below if you want to try it." },
"background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } }
]
}'Can the prompt set tone, gestures, or camera moves?
No. In current code the talking-video create has no field for tone, gestures, or camera moves, and it refuses any field outside its short list with a 400. Keep script to spoken words too: in current code each clip asks for exactly the dialogue you wrote, so a stage direction such as "(laughs)" goes in as words to say; how many words fit in 60 seconds covers the other script rules.
For tone and pace, speak the lines with TTS 1.0 and lip sync the avatar with Fabric, as in AI avatar emotion and speaking speed. For set gestures, see AI avatar hand gestures.
Why not write one text-to-video prompt with a voice-over?
Because the mouth won't match. Sume's model guide says video models don't lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath. Talking shots come from the avatar route, or from a still lip-synced to TTS with VEED Fabric 1.0; the talking head video API comparison weighs the routes. Wordless B-roll can still come from a video prompt.
What are the limits?
- The whole plan must estimate at 4–60 seconds, and
video_inputstakes up to 20 scenes. - One avatar and one shared scene per video.
- English speech only, in current code: the clip prompt asks for English dialogue.
- Media fields must be public HTTPS URLs. A request inside the time window can still fail with
content_policy_rejectedwhen a provider rejects the generated content.
Sources
Related posts
More in Sume Avatar 1.0
- What is a talking head video? Meaning and how to make one
A talking head video shows one person speaking to the camera, framed from about the chest up. What it means, what it's for, and how AI makes one.
- What is AI UGC? Meaning, how it's made, and how it differs
AI UGC is UGC-style video made with AI: a generated presenter speaks your script to camera, the way a creator would on a phone. How it differs.
- What is an AI avatar? What it does and what it can't
An AI avatar is a digital presenter, a face and a voice, that speaks the script you write in a video. What it does, how it's made, and its limits.
- What is an AI digital twin? Twins vs photo and stock avatars
An AI digital twin is a custom avatar of one real person that speaks new scripts in their likeness. HeyGen trains it on footage; Synthesia starts from a photo.
Written by Sume