Avatar explainer: product_image or a photo scene, what each does
In Sume Avatar 1.0, product_image adds $0.01, $0.013 or $0.03 a second by tier; a photo scene sets the background. When to send which for a product explainer.
In Sume Avatar 1.0, product_image tells the video what the product is and adds $0.01, $0.013 or $0.03 per second on standard, plus and max. A photo scene, sent as scene with type photo and an image_url, sets where the avatar stands. For a product explainer, send product_image when the product must appear, and a scene when the location matters.
The Avatar Video docs (read 2026-10-06) list both as optional inputs. Omit product_image for a productless video. Media fields must be fetchable public HTTPS URLs.
What each field is for
The product image is a subject. The scene is a place. Mixing them up is the common mistake: a store-shelf photo sent as product_image asks the avatar to present the shelf, not to stand in front of it.
| Field | Shape | What it controls | Rate effect |
|---|---|---|---|
| product_image | Public HTTPS image URL | The product the avatar presents | Plus $0.01, $0.013 or $0.03 a second |
| scene (photo) | type photo and image_url | The location reference | Pricing table has no separate column |
| scene (prompt) | type prompt and a prompt | Scene direction in words | Pricing table has no separate column |
The cost of the product image
Over a 30-second explainer the product image adds $0.30 on standard, $0.39 on plus and $0.90 on max. That is small next to the clip itself, so the question is accuracy, not cost: does the product have to be recognizable? If it is a physical item with a logo or a specific shape, send it. If the explainer is about a service or an idea, skip it and keep the rate lower.
The rate table has two columns, with a product image and without. How a scene-only request is priced follows from whether product_image is present, so read the estimate on a preview if you send a scene alone.
Check before the full render
Create an Avatar video preview with both fields and look at the first frame. For multi-scene plans the preview returns one still per scene, and later stills are pose-anchored continuations of the first. If the product is tiny, hidden or wrongly sized, change the inputs and make a new preview; the docs say structural changes such as the scene need a new one.
Keep expectations honest. A generated frame can change small details of a product, so do not use a render as proof of a label, a size or a claim. Put those in the captions and the copy.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar script over 60 seconds: split it into two Avatar 1.0 jobs
Sume accepts avatar scripts that it estimates at 4 to 60 seconds. Longer scripts need to be shortened or split into jobs; here is how to split without a jump.
- Avatar video for a 4:5 feed: render 3:4, then crop to 0.9375 height
Sume Avatar 1.0 has no 4:5 ratio. Render 3:4, crop the height to 0.9375 with video filter for $0.02, and caption after the crop so nothing is cut off.
- Avatar video ratio: five options for LinkedIn, Shorts and feeds
Sume Avatar 1.0 renders 1:1, 3:4, 9:16, 4:3 and 16:9 at 720p. Which ratio to pick for LinkedIn, Shorts, a feed post or a landing page, and what changes in cost.
- Bakery pre-order deadline: a 20-second avatar clip with the cake photo
A 20-second pre-order reminder with the cake photo as product_image costs about $3.88 on Standard, $5.16 on Plus or $11.60 on Max at Sume's rates.
Written by Sume