Avatar explainer: product_image or a photo scene, what each does

In Sume Avatar 1.0, product_image adds $0.01, $0.013 or $0.03 a second by tier; a photo scene sets the background. When to send which for a product explainer.

4 min readSume
All posts

In Sume Avatar 1.0, product_image tells the video what the product is and adds $0.01, $0.013 or $0.03 per second on standard, plus and max. A photo scene, sent as scene with type photo and an image_url, sets where the avatar stands. For a product explainer, send product_image when the product must appear, and a scene when the location matters.

The Avatar Video docs (read 2026-10-06) list both as optional inputs. Omit product_image for a productless video. Media fields must be fetchable public HTTPS URLs.

What each field is for

The product image is a subject. The scene is a place. Mixing them up is the common mistake: a store-shelf photo sent as product_image asks the avatar to present the shelf, not to stand in front of it.

Optional avatar video inputs (read 2026-10-06)
FieldShapeWhat it controlsRate effect
product_imagePublic HTTPS image URLThe product the avatar presentsPlus $0.01, $0.013 or $0.03 a second
scene (photo)type photo and image_urlThe location referencePricing table has no separate column
scene (prompt)type prompt and a promptScene direction in wordsPricing table has no separate column

The cost of the product image

Over a 30-second explainer the product image adds $0.30 on standard, $0.39 on plus and $0.90 on max. That is small next to the clip itself, so the question is accuracy, not cost: does the product have to be recognizable? If it is a physical item with a logo or a specific shape, send it. If the explainer is about a service or an idea, skip it and keep the rate lower.

The rate table has two columns, with a product image and without. How a scene-only request is priced follows from whether product_image is present, so read the estimate on a preview if you send a scene alone.

Check before the full render

Create an Avatar video preview with both fields and look at the first frame. For multi-scene plans the preview returns one still per scene, and later stills are pose-anchored continuations of the first. If the product is tiny, hidden or wrongly sized, change the inputs and make a new preview; the docs say structural changes such as the scene need a new one.

Keep expectations honest. A generated frame can change small details of a product, so do not use a render as proof of a label, a size or a claim. Put those in the captions and the copy.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume