How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
An AI UGC ad looks less staged when the first frame looks like a phone video: chest-up framing, ordinary light, a real-looking room, and words people say out loud. On Sume you control that through the scene input, the script, and an approved preview, not through a prompt that says make it look real.
Sume's first-frame stage is already written toward the phone look: a fixed phone camera, close-to-medium chest-up framing, uneven everyday light, no studio polish. What you add are the choices below, read from the avatar video docs and preview docs on 2026-10-03. Whether the result reads as real is a judgment call, so the last section is a review step.
Which input controls the look of the frame?
The docs' own multi-scene example uses the prompt "Casual bedroom framing, native UGC lighting". Media fields must be public HTTPS URLs.
| Input | What it does | UGC use |
|---|---|---|
| scene: { type: prompt } | Directs the setting in words | Name a place and a light: kitchen counter, daytime window |
| scene: { type: photo, image_url } | Uses your photo as the scene reference | A real room, desk or shop floor |
| product_image | Puts your product in the frame | Adds $0.010-$0.030 per second by tier |
| aspect_ratio | 1:1, 3:4, 9:16, 4:3, 16:9 | 9:16 is the default |
| quality | standard, plus, max | Judge the look on the preview, not the tier |
How should the script read?
- Write for speech, not print: short sentences, contractions, one idea per line.
- Open with the reaction, as the docs' sample does: "Wait, this turned one selfie into a whole video?"
- Keep each spoken scene under Sume's 2,000 characters, and the whole video inside 4 to 60 seconds.
- Use a
silencescene for a beat of non-speech, such as a product shot, instead of padding with filler words. - Add captions with the
captionsfield; preview stills are never captioned.
Why approve a preview first?
The preview renders first-frame stills without the full render, so you can check framing, the room and the product before you pay for motion. A multi-scene preview gives one still per scene, and later stills are pose-anchored continuations of the first. If the frame is off, regenerate the stills; the script and avatar stay fixed. When the frame looks right, call generate-video on the preview id, and optionally override quality for the final only, since preview stills are tier-independent.
Changing the script, scene, avatar or aspect ratio needs a new preview, because those are structural.
What does a request look like?
curl -X POST https://api.sume.com/v1/avatar-video-previews \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ugc-preview-001" \
-d '{
"avatar_handle": "sume_clawra",
"aspect_ratio": "9:16",
"quality": "standard",
"scene": {"type": "prompt", "prompt": "Casual kitchen counter, daytime window light"},
"script": "Okay, I did not expect this to work. Watch what happens when I drop in one photo."
}'What should I review before the final render?
Read every still as a viewer scrolling a feed would. Is the face the one you chose, and does it stay the same face across scene stills? Are the hands and product plausible? Does the room look like a place someone would actually record in? If a still fails, regenerate; the preview keeps the stored request, so only the stills change. For a long video, you can also sample frames after the render with video inspect to catch drift.
Run the cheap tier while you search for the right frame and the right words, then override quality on the approved preview for the final render. That keeps the spend on the version you will publish. Disclosure is the last check: an avatar clip still needs the AI label your platform or market requires.
What belongs in the script?
- Spoken sentences: write for the ear, with contractions and short lines.
- One idea per scene: a multi-scene video takes up to 20 scenes, each up to 2000 characters, but a UGC ad rarely needs more than a few.
- A beat of reaction: let the first line sound like a thought, not a headline.
- No stage directions inside the spoken text: anything you put in the script is read aloud.
What can I not control?
Avatar 1.0 speaks English only in code, and each video has one avatar in one shared scene, so you cannot cut between locations inside a single video. For a multi-location ad, render separate videos and join them in your editor. Previews let you check the look first, then generate-video renders the one you approve.
What can you not change?
One video holds one avatar and one shared scene, so a mid-video location change means joining clips. Resolution is 720p, and Avatar 1.0 is English-only in code. Treat the preview as the review step: read each still for the face, hands and product, and regenerate rather than ship a frame you would not post yourself.
Sources
Related posts
- Avatar video previews: approve the first frame before rendering
- Avatar video preview: approve the first frame, then pick quality
- Can one avatar UGC ad change location mid-video? One scene only
- Hook variations for UGC ads: swap the hook, reuse the body
- AI avatar prompts: how to describe a talking presenter
More in Sume Avatar 1.0
- Shortest AI avatar video: 4 seconds, from $0.74 on Sume Standard
Sume avatar videos run 4 to 60 seconds. Per-second rates for Standard, Plus and Max, the 4 second floor, and what 15, 30 and 60 second clips cost.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume