UGC-style avatar video: a 9:16 casual scene prompt that works on Sume
How to ask Sume Avatar 1.0 for a phone-filmed look: 9:16, a casual background prompt, silence beats, product image. Request body and what it costs per tier.
To get a UGC-style look from Sume Avatar 1.0, request aspect_ratio: "9:16" (the default), write a casual background prompt such as "Casual bedroom framing, native UGC lighting", and structure the script as short scenes with video_inputs. The output is 720p, which is the only resolution at this time. Sume avatars are async jobs: you submit a request, get a job id back, and read the finished video later. There is no live video session.
The scene prompt is the part people get wrong. It should describe the framing and the light, not a story. This post shows a three-scene request lifted from the pattern in the Generate avatar video docs, then prices it.
The three-beat structure
A UGC clip usually has a hook, a demo, and a call to action. In video_inputs, each scene has an id, a voice, and a background. A spoken scene uses voice.type: "text" with one of script or input_text, and a duration. A silence scene uses voice.type: "silence" with a duration and no text, which suits the beat where the product is shown.
The total planned duration must still fit in the 4 to 60 second window. Sume's current execution expects the scene backgrounds to resolve to one shared scene, so repeat the same background prompt in each scene as the docs example does.
| Scene id | Voice type | Seconds | What happens |
|---|---|---|---|
| hook | text | 3 | Short surprised line to camera |
| demo | silence | 4 | Product shown, no speech |
| cta | text | 5 | One sentence telling the viewer what to do |
The request
Use product_image if the item must appear. Add captions if the clip plays muted; the default caption style is slam.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ugc-bottle-001" \
-d '{
"avatar_handle": "product_host",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "Wait, this bottle kept my coffee hot all day?", "duration": 3},
"background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}},
{"id": "demo", "voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}},
{"id": "cta", "voice": {"type": "text", "script": "Grab one from the link below.", "duration": 5},
"background": {"type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting"}}
]
}'Cost of the 12-second clip
The silence beat counts toward the planned duration, so the clip bills as 12 seconds.
| Tier | No product | With product |
|---|---|---|
| standard | $2.208 | $2.328 |
| plus | $2.94 | $3.096 |
| max | $6.60 | $6.96 |
Be straight with viewers
A UGC-style look does not make the person real. The avatar is generated. Do not present it as a customer, do not invent a testimonial, and label the video as synthetic where the platform or the law asks for it. A casual look is a style choice, and the label is a separate decision that you should make for every ad.
Writing the scene prompt
Keep the background prompt short and physical: where the camera is, what the light is like, how casual the framing is. "Casual bedroom framing, native UGC lighting" works because it names a place and a lighting style and says nothing about the person. Avoid asking the background to carry a plot, a logo, or text, because the clip is built around the avatar and the scene is only the setting.
If you have a real photograph of the setting, use scene: {"type": "photo", "image_url": "..."} in a single-script request instead of a prompt. Both are optional, both must be public HTTPS URLs when they are images, and Sume checks the URL before it submits the render.
For several variants of one ad, change the hook line and keep everything else. The cheap way to compare hooks is to render each at standard and send only the winner through max. Since each variant is its own job, you can run them in parallel within your workspace's concurrency limits.
Common mistakes
The plus tier is the default when you omit quality, and it is the balanced path. standard is the fastest Sume execution path, which makes it the right tier for hook tests. max is the highest quality and has slower turnaround, which makes it the tier for the version that will run as a paid ad. Because UGC-style ads live or die on the first three seconds, spend your testing budget on hooks at standard, and spend the final render budget on the winner.
- A silence scene with a
scriptorinput_textfield. Silence takes only aduration. - A single request with both
scriptandvideo_inputs. Provide one. - Scenes whose backgrounds do not resolve to one shared scene.
- A total planned duration under 4 seconds or over 60.
- Captions in
slamstyle on a Korean script; Sume answers400 caption_hangul_text_latin_styleand wants a Hangul style.
Sources
Related posts
More in Use cases
- Virtual staging an empty room photo with AI: GPT Image 2.5 via API
Stage an empty room photo with openai/gpt-image-2.5: the room plus up to 15 furniture references, an optional mask, and about $0.08 per staged image at high.
- How do I make a winter tire changeover promo video for my garage?
Three 6-second Wan 3.0 clips, a 420-character voiceover, a duck-under music bed and booking captions for a garage: $2.695 on Sume at 720p.
- Write a lip-sync script in 5 to 14.8 second beats (H3 Max voice-over)
MiniMax H3 Max lip-sync on Sume takes audio of 5 to 14.8 seconds. Write the script in beats that fit the window, then voice each beat; the price per beat.
- YouTube bumper ad 5-6 seconds: generate at 6 or trim on Sume
YouTube's help page puts bumper ads at 5 to 6 seconds. Which Sume video models can generate a 6 second clip directly, and when a $0.02 trim is the better route.
Written by Sume