How to make an AI UGC ad look less staged with Avatar 1.0

Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.

5 min readSume
All posts

An AI UGC ad looks less staged when the first frame looks like a phone video: chest-up framing, ordinary light, a real-looking room, and words people say out loud. On Sume you control that through the scene input, the script, and an approved preview, not through a prompt that says make it look real.

Sume's first-frame stage is already written toward the phone look: a fixed phone camera, close-to-medium chest-up framing, uneven everyday light, no studio polish. What you add are the choices below, read from the avatar video docs and preview docs on 2026-10-03. Whether the result reads as real is a judgment call, so the last section is a review step.

Which input controls the look of the frame?

The docs' own multi-scene example uses the prompt "Casual bedroom framing, native UGC lighting". Media fields must be public HTTPS URLs.

Avatar video inputs that change the look of a UGC ad, read 2026-10-03
InputWhat it doesUGC use
scene: { type: prompt }Directs the setting in wordsName a place and a light: kitchen counter, daytime window
scene: { type: photo, image_url }Uses your photo as the scene referenceA real room, desk or shop floor
product_imagePuts your product in the frameAdds $0.010-$0.030 per second by tier
aspect_ratio1:1, 3:4, 9:16, 4:3, 16:99:16 is the default
qualitystandard, plus, maxJudge the look on the preview, not the tier

How should the script read?

  • Write for speech, not print: short sentences, contractions, one idea per line.
  • Open with the reaction, as the docs' sample does: "Wait, this turned one selfie into a whole video?"
  • Keep each spoken scene under Sume's 2,000 characters, and the whole video inside 4 to 60 seconds.
  • Use a silence scene for a beat of non-speech, such as a product shot, instead of padding with filler words.
  • Add captions with the captions field; preview stills are never captioned.

Why approve a preview first?

The preview renders first-frame stills without the full render, so you can check framing, the room and the product before you pay for motion. A multi-scene preview gives one still per scene, and later stills are pose-anchored continuations of the first. If the frame is off, regenerate the stills; the script and avatar stay fixed. When the frame looks right, call generate-video on the preview id, and optionally override quality for the final only, since preview stills are tier-independent.

Changing the script, scene, avatar or aspect ratio needs a new preview, because those are structural.

What does a request look like?

curl -X POST https://api.sume.com/v1/avatar-video-previews \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ugc-preview-001" \
  -d '{
    "avatar_handle": "sume_clawra",
    "aspect_ratio": "9:16",
    "quality": "standard",
    "scene": {"type": "prompt", "prompt": "Casual kitchen counter, daytime window light"},
    "script": "Okay, I did not expect this to work. Watch what happens when I drop in one photo."
  }'

What should I review before the final render?

Read every still as a viewer scrolling a feed would. Is the face the one you chose, and does it stay the same face across scene stills? Are the hands and product plausible? Does the room look like a place someone would actually record in? If a still fails, regenerate; the preview keeps the stored request, so only the stills change. For a long video, you can also sample frames after the render with video inspect to catch drift.

Run the cheap tier while you search for the right frame and the right words, then override quality on the approved preview for the final render. That keeps the spend on the version you will publish. Disclosure is the last check: an avatar clip still needs the AI label your platform or market requires.

What belongs in the script?

  • Spoken sentences: write for the ear, with contractions and short lines.
  • One idea per scene: a multi-scene video takes up to 20 scenes, each up to 2000 characters, but a UGC ad rarely needs more than a few.
  • A beat of reaction: let the first line sound like a thought, not a headline.
  • No stage directions inside the spoken text: anything you put in the script is read aloud.

What can I not control?

Avatar 1.0 speaks English only in code, and each video has one avatar in one shared scene, so you cannot cut between locations inside a single video. For a multi-location ad, render separate videos and join them in your editor. Previews let you check the look first, then generate-video renders the one you approve.

What can you not change?

One video holds one avatar and one shared scene, so a mid-video location change means joining clips. Resolution is 720p, and Avatar 1.0 is English-only in code. Treat the preview as the review step: read each still for the face, hands and product, and regenerate rather than ship a frame you would not post yourself.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume