UGC avatar scene prompt: why dim, glary lighting words are kept
Sume's first-frame step copies lighting words from your scene prompt as written, so ordinary light like TV glow or a side window stays. How to write it.
If you want a UGC-style avatar ad that looks like a phone video and not a studio spot, write the lighting you actually want in scene.prompt and keep it plain. Sume's first-frame step is built to copy the lighting words from your scene prompt into the image prompt without upgrading them, so "TV glow" or "overhead light" stays TV glow or overhead light.
That behaviour comes from the instruction text that the Avatar 1.0 workflow sends to its first-frame planner, read in the repository on 2026-10-11. The public contract is on the Generate avatar video page: scene: { "type": "prompt", "prompt": "..." } gives scene direction, and scene: { "type": "photo", "image_url": "..." } gives a photo reference. Below is what the prompt step does with your words and how to use it.
What does the first-frame step do with my scene words?
Before the talking video renders, Sume builds a first frame from the avatar image plus your scene. The planner instruction tells the model to act as a literal assembler: the scene is the authority for background and lighting, lighting nouns and adjectives are copied across "without synonym substitution", and it must not make the scene prettier, cleaner, softer, brighter, warmer or more cinematic.
The same instruction asks for the frame to feel like a single imperfect freeze from the first second of a real talking video: a mid-speech mouth shape or a blink, relaxed shoulders, mild compression, slight wide-angle proximity and uneven exposure. It avoids symmetrical influencer poses, perfect smile portraits and cinematic lighting. In short, the default is already tilted toward phone-video realism, and your scene prompt is the lever on top of it.
Which lighting words does it keep as written?
The instruction names the kinds of ordinary light it should keep rather than fix. Use that list as a vocabulary. The table maps each rule to what you should write.
| If you want | Write in scene.prompt | Rule that applies |
|---|---|---|
| Screen light on the face | mixed monitor glow, TV glow | Kept: ordinary lighting is not cleaned up |
| A lived-in room | lamp light, side window light, dim interior | Kept: no invented extra light sources |
| Flat daytime look | overcast daylight, overhead light | Kept: lighting nouns copied unchanged |
| Imperfect exposure | glare, uneven exposure | Kept: ugly lighting stays if requested |
| A specific backdrop | name only the objects you want | No background objects are added that you did not ask for |
What does it refuse to add?
The planner is told not to add bokeh, soft glow, beauty lighting, cinematic colour grading, studio polish or flattering illumination. It also should not add a second light source you did not mention. That is useful when you compare variants: if two scene prompts differ only in one lighting word, the first frames should differ in that one respect, rather than both drifting toward the same glossy look.
The flip side is that vague words get no help. "Nice lighting" gives the planner nothing concrete to copy, so name the source and the time of day instead, for example "side window light, late afternoon, one lamp on in the corner".
Does the product follow the same literal rule?
Yes, in a stricter form. If you send product_image, the instruction treats it as the exact product reference and asks for one product only, not duplicated, redesigned, hidden or genericized. If you omit it, the frame should have no focal product, package or intentional brand placement; objects from your scene prompt may appear only as secondary background or foreground details.
So a scene prompt that says "on a kitchen counter" can put a kettle in the background, but it does not turn the kettle into the thing being sold. To feature a product, pass the image. Pricing differs slightly with a product image; the per-second rates are on the price ladder.
How do I test a scene prompt before the full render?
Create an Avatar video preview. It runs only the first-frame stage and returns preview_image_url, so you can read the lighting before you pay for a full talking render. If the frame is off, call regenerate on the same preview, which reuses the stored request and refreshes the stills. If you want to change the scene prompt itself, that is a structural change and needs a new preview.
When the still is right, call generate-video on the preview id. Sume reuses the approved first frame, and you can still choose the final tier at that step. For wording examples on the other fields, see a casual UGC scene prompt.
Sources
Related posts
More in Sume Avatar 1.0
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
- Multi-scene avatar video API: build one video from ordered scenes
Send ordered video_inputs instead of one script to compose spoken and silent scenes into one 4-60 second avatar video. Scene fields, rules, and limits.
Written by Sume