Same set every episode: an Avatar 1.0 scene photo or prompt background
Avatar Video 1.0 scenes are a prompt or a photo. Use one photo for the set in every episode, or reuse one prompt word for word, to keep the same place.
Use the same photo scene in every episode. In Avatar Video 1.0 a scene is either {type: prompt} or {type: photo, image_url}, and the photo form puts the presenter in the place your image shows. Keep that image at one public HTTPS URL in your season config so every episode starts from the same set.
The two scene types
| Scene | Field | Best for |
|---|---|---|
| Prompt | {type: "prompt"} with a text description | Exploring a look, a one-off episode |
| Photo | {type: "photo", image_url}, a public HTTPS image | A recurring set, a real office or studio |
| Silence beat | voice.type: "silence" with a required duration, no script | A pause between lines in multi-scene videos |
Why a photo wins for a series
A prompt is interpreted again every time, so the room varies a little from one episode to the next. A photo is a fixed input. If your set is a real desk, take one photo of it, keep the lighting the same on shoot days, and use that file for every episode. Change the photo on purpose when the story moves somewhere else.
Keep it stable
- Host the image at a durable HTTPS address, not a temporary link. Sume rejects localhost and private-network URLs.
- Leave space in the photo where the presenter will appear; a cluttered centre fights with them.
- Only one resolved avatar is supported per final video, so a two-presenter episode needs two videos cut together on a timeline.
- Preview first with
POST /v1/avatar-video-previewsand regenerate before you pay for the full video.
When to use a prompt
A prompt scene is quicker when the set changes every episode, such as a travel series. Write the scene once as a sentence, hold it in your season config and reuse the exact words. Small wording changes give a different room, so edit the text only on purpose.
Both types work in multi-scene video_inputs, so you can open an episode on a photo of the set and cut to a prompt scene for a flashback, as long as the same avatar appears throughout.
Write down, once, which scene type each recurring location uses and keep that table in your repo. In practice a series has two or three settings, such as an office, a kitchen and a street, and each one maps to a single photo URL or a single sentence. When an editor asks why episode 8 looks different, the answer is then one line in the config.
Check one more thing before a season: the photo's orientation should match the aspect ratio you request. A landscape photo in a 9:16 video crops, and what survives the crop is a choice you want to preview, not discover.
Sources
Related posts
More in Sume Avatar 1.0
- Sponsored episode with an Avatar 1.0 presenter: product image cost
Adding a product image to Avatar Video 1.0 costs $0.01 to $0.03 more per second. Standard $0.184 to $0.194, Plus $0.245 to $0.258, Max $0.55 to $0.58.
- Sume Avatar max costs 2.24x plus and 3x standard: when to pay it
Sume Avatar 1.0 is $0.184 a second on standard, $0.245 plus, $0.55 max: ratios 1 : 1.33 : 2.99. Iterate on standard, ship plus, pay max for hero ads.
- AI talking avatar video cost per minute: Sume tiers
Sume avatar video is $0.184, $0.245 or $0.55 per second by tier: $11.04, $14.70 or $33.00 a minute. 100 thirty-second clips cost $552 to $1,650.
- Griffin-Lite study: 81% confident about the AI, 79% about people
Tavus reports participants were 81% confident about the AI and 79% about real people. Confidence did not track the truth, so label AI avatar clips yourself.
Written by Sume