Avatar photo scene: a talking host in your showroom for year-end
Use a scene photo reference in Sume Avatar 1.0 so a talking host stands in your own showroom or shop for a year-end sale clip.
Scene as photo, scene as prompt
The avatar-video request accepts scene: { "type": "photo", "image_url": "https://..." } for a photo scene reference, or scene: { "type": "prompt", "prompt": "..." } for text direction. For a year-end sale in a real space, use the photo form with a picture of the showroom, shop floor or counter.
Media fields must be fetchable public HTTPS URLs. Localhost, private-network and non-HTTPS URLs are rejected before generation.
A single-script request
Omit product_image for a productless clip, or pass a product image to feature one item. Use script for a single speaker; the estimated duration must be 4 to 60 seconds, so shorten longer scripts or split them into multiple jobs.
curl -sS -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: showroom-yearend-001" \
-d '{
"avatar_handle": "sume_clawra",
"script": "Year-end clearance starts today. Floor models are marked down until the thirty-first.",
"scene": { "type": "photo", "image_url": "https://cdn.example.com/showroom.jpg" },
"aspect_ratio": "9:16",
"quality": "plus"
}'Choose the photo for the shot
Because the avatar is placed into the scene, choose a photo with clear floor space and even light. Use one photo per clip; the current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene.
| Input | Form | Effect |
|---|---|---|
| product_image | public HTTPS URL; optional | Featured product; omit for a productless clip |
| scene.type = prompt | text direction | Describes the setting |
| scene.type = photo | public HTTPS image_url | Photo scene reference |
When the location changes
A clip that needs two locations is two jobs joined afterwards, since one request resolves to one shared scene. Generate each clip with its own scene, then assemble them with a Sume timeline route. Keep the same avatar_handle so the host stays the same person.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar script over 60 seconds: split it into jobs and join the clips
Sume Avatar 1.0 accepts 4 to 60 seconds per job. For a longer explainer, split the script at sentence breaks, send one job per part and keep the order.
- Write an avatar script that sounds authentic in 4 to 60 seconds
HeyGen says 64.6% trust avatars that sound authentic. Draft a Sume avatar script in the 4-60 second window with scene beats and a silence gap.
- Avatar video aspect ratios: 9:16, 1:1, 4:3 or 16:9 for ads?
Sume avatar videos render in 1:1, 3:4, 9:16, 4:3 or 16:9 at 720p. Which ratio to pick for feed, story and listing placements, and what the default is.
- Make an avatar video look professional: scene photo of your office
53.9% of shelved videos were held back for not looking professional enough, per HeyGen. Set a Sume avatar scene from your own photo, prompt or per-scene image.
Written by Sume