UGC-style avatar video: a 9:16 casual scene prompt that works on Sume

How to ask Sume Avatar 1.0 for a phone-filmed look: 9:16, a casual background prompt, silence beats, product image. Request body and what it costs per tier.

4 min readSume
All posts

To get a UGC-style look from Sume Avatar 1.0, request aspect_ratio: "9:16" (the default), write a casual background prompt such as "Casual bedroom framing, native UGC lighting", and structure the script as short scenes with video_inputs. The output is 720p, which is the only resolution at this time. Sume avatars are async jobs: you submit a request, get a job id back, and read the finished video later. There is no live video session.

The scene prompt is the part people get wrong. It should describe the framing and the light, not a story. This post shows a three-scene request lifted from the pattern in the Generate avatar video docs, then prices it.

The three-beat structure

A UGC clip usually has a hook, a demo, and a call to action. In video_inputs, each scene has an id, a voice, and a background. A spoken scene uses voice.type: "text" with one of script or input_text, and a duration. A silence scene uses voice.type: "silence" with a duration and no text, which suits the beat where the product is shown.

The total planned duration must still fit in the 4 to 60 second window. Sume's current execution expects the scene backgrounds to resolve to one shared scene, so repeat the same background prompt in each scene as the docs example does.

Scene plan for a 12-second UGC clip (Sume docs, read 2026-10-07)
Scene idVoice typeSecondsWhat happens
hooktext3Short surprised line to camera
demosilence4Product shown, no speech
ctatext5One sentence telling the viewer what to do

The request

Use product_image if the item must appear. Add captions if the clip plays muted; the default caption style is slam.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ugc-bottle-001" \
  -d '{
    "avatar_handle": "product_host",
    "aspect_ratio": "9:16",
    "quality": "plus",
    "video_inputs": [
      {"id": "hook", "voice": {"type": "text", "script": "Wait, this bottle kept my coffee hot all day?", "duration": 3},
       "background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}},
      {"id": "demo", "voice": {"type": "silence", "duration": 4},
       "background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}},
      {"id": "cta", "voice": {"type": "text", "script": "Grab one from the link below.", "duration": 5},
       "background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}}
    ]
  }'

Cost of the 12-second clip

The silence beat counts toward the planned duration, so the clip bills as 12 seconds.

12-second UGC-style clip by tier, with and without a product image (provider-pricing, read 2026-10-07)
TierNo productWith product
standard$2.208$2.328
plus$2.94$3.096
max$6.60$6.96

Be straight with viewers

A UGC-style look does not make the person real. The avatar is generated. Do not present it as a customer, do not invent a testimonial, and label the video as synthetic where the platform or the law asks for it. A casual look is a style choice, and the label is a separate decision that you should make for every ad.

Writing the scene prompt

Keep the background prompt short and physical: where the camera is, what the light is like, how casual the framing is. "Casual bedroom framing, native UGC lighting" works because it names a place and a lighting style and says nothing about the person. Avoid asking the background to carry a plot, a logo, or text, because the clip is built around the avatar and the scene is only the setting.

If you have a real photograph of the setting, use scene: {"type": "photo", "image_url": "..."} in a single-script request instead of a prompt. Both are optional, both must be public HTTPS URLs when they are images, and Sume checks the URL before it submits the render.

For several variants of one ad, change the hook line and keep everything else. The cheap way to compare hooks is to render each at standard and send only the winner through max. Since each variant is its own job, you can run them in parallel within your workspace's concurrency limits.

Common mistakes

The plus tier is the default when you omit quality, and it is the balanced path. standard is the fastest Sume execution path, which makes it the right tier for hook tests. max is the highest quality and has slower turnaround, which makes it the tier for the version that will run as a paid ad. Because UGC-style ads live or die on the first three seconds, spend your testing budget on hooks at standard, and spend the final render budget on the winner.

  • A silence scene with a script or input_text field. Silence takes only a duration.
  • A single request with both script and video_inputs. Provide one.
  • Scenes whose backgrounds do not resolve to one shared scene.
  • A total planned duration under 4 seconds or over 60.
  • Captions in slam style on a Korean script; Sume answers 400 caption_hangul_text_latin_style and wants a Hangul style.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume