UGC video prompt for AI avatars: what goes in each field

A UGC video prompt for an AI avatar is three inputs, not one: the creator's words, a scene prompt for the setting and light, and a product image.

5 min readSume
All posts

A UGC video prompt for an AI avatar isn't one block of text. A talking UGC clip has three parts that go in different places: the words the creator says, the setting and light, and the product. In Sume's Avatar 1.0 API, the words go in script (or each scene's voice), the setting in a scene prompt or photo, and the product in product_image. The person's look was fixed when you created the avatar.

Field rules come from Sume's Generate avatar video docs and the Sume API reference, read on 2026-09-28; anything called current behavior is read from Sume's code. For the lines themselves, see UGC ad script for AI avatar videos; for the person, how to write an AI avatar prompt.

What goes in each part of a UGC avatar prompt?

Map each thing you would type into a one-box generator to its field:

From Generate avatar video and the Sume API reference, read 2026-09-28.
Part of the clipFieldWhat to write
The personavatar_handleNothing: the look comes from the avatar you created
The wordsscript, or a scene's voice with type: "text"Only the lines the creator says
A pause for a demoA scene's voice with type: "silence" and a durationNo words; the seconds it lasts
Setting and lightscene (a prompt or a photo), or each scene's backgroundWhere it's filmed and how it's lit
The productproduct_imageA public HTTPS image URL; leave it out for no product
The frameaspect_ratio9:16 by default; also 1:1, 3:4, 4:3, 16:9

How do I write the scene prompt?

Name the place, the light, and the framing in one short phrase. The docs' own UGC example uses "Casual bedroom framing, native UGC lighting". Other phrases in that style:

  • Kitchen counter, morning window light, close framing.
  • Parked car interior, daylight through the windshield.
  • Bathroom mirror, warm vanity light, chest-up framing.
  • Use one setting per video: in current code, scenes with different backgrounds are refused with a 400, so repeat the same prompt on each scene. A real room works too, as a public HTTPS photo in scene.

What does a full UGC avatar request look like?

A spoken hook, a silent demo beat, and a spoken call to action, each with the same background, plus the product:

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ugc-serum-v1" \
  -d '{
    "avatar_handle": "ugc_creator",
    "product_image": "https://example.com/serum.png",
    "video_inputs": [
      { "id": "hook",
        "voice": { "type": "text", "script": "Okay, this is the serum everyone keeps asking me about." },
        "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
      { "id": "demo",
        "voice": { "type": "silence", "duration": 4 },
        "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } },
      { "id": "cta",
        "voice": { "type": "text", "script": "The link is below if you want to try it." },
        "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } }
    ]
  }'

Can the prompt set tone, gestures, or camera moves?

No. In current code the talking-video create has no field for tone, gestures, or camera moves, and it refuses any field outside its short list with a 400. Keep script to spoken words too: in current code each clip asks for exactly the dialogue you wrote, so a stage direction such as "(laughs)" goes in as words to say; how many words fit in 60 seconds covers the other script rules.

For tone and pace, speak the lines with TTS 1.0 and lip sync the avatar with Fabric, as in AI avatar emotion and speaking speed. For set gestures, see AI avatar hand gestures.

Why not write one text-to-video prompt with a voice-over?

Because the mouth won't match. Sume's model guide says video models don't lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath. Talking shots come from the avatar route, or from a still lip-synced to TTS with VEED Fabric 1.0; the talking head video API comparison weighs the routes. Wordless B-roll can still come from a video prompt.

What are the limits?

  • The whole plan must estimate at 4–60 seconds, and video_inputs takes up to 20 scenes.
  • One avatar and one shared scene per video.
  • English speech only, in current code: the clip prompt asks for English dialogue.
  • Media fields must be public HTTPS URLs. A request inside the time window can still fail with content_policy_rejected when a provider rejects the generated content.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume