Map avatar video scene ids to onboarding steps with scene_previews
Give each video_inputs scene a stable id and Sume returns it in scene_previews with start, end and duration, so one clip can drive per-step chapters.
Set video_inputs[].id to your own stable step key, such as step-1-connect-account, and Sume echoes it back in scene_previews[].id along with start_time_seconds, end_time_seconds, and duration_seconds. That gives you chapter markers and per-step stills for one onboarding clip without keeping your own bookkeeping of which sentence landed where.
What the API accepts and returns
In a multi-scene request, each video_inputs item takes an optional id (1-100 characters, described as a client-provided stable scene id for metadata) and a required voice. The array holds 1-20 scenes, and each spoken scene has up to 2,000 characters of script or input_text. The total planned duration must still fall in the 4-60 second window. A second per-scene id use, scene_previews[].id, is documented as the optional caller-provided id when available from video_inputs[].
| Field | Where | Use |
|---|---|---|
| video_inputs[].id | Request | Your step key, 1-100 characters |
| scene_previews[].index | Response | Position of the scene |
| scene_previews[].id | Response | Your id when available |
| scene_previews[].start_time_seconds | Response | Chapter start |
| scene_previews[].duration_seconds | Response | Scene length |
| scene_previews[].preview_image_url | Response | First-frame still for the scene |
Build chapters from the response
Rather than a separate clip for each step, you can render one short onboarding video with several steps and still link into it by time. This works only up to the 60-second ceiling, so a long checklist should be split by topic into several jobs; that limit and the split strategy are covered by the one-job-per-step approach in the related posts.
Keep a lookup of your step keys. When the resource is ready, read scene_previews, drop entries without an id, and write a chapter list for your player.
def chapters(scene_previews):
out = []
for s in scene_previews or []:
if not s.get("id"):
continue
out.append({
"step": s["id"],
"start": s.get("start_time_seconds"),
"still": s.get("preview_image_url"),
})
return out
sample = [
{"index": 0, "id": "step-1", "start_time_seconds": 0,
"preview_image_url": "https://media.sume.com/a.png"},
{"index": 1, "id": "step-2", "start_time_seconds": 9.5,
"preview_image_url": None},
]
print(chapters(sample))Checks before you rely on it
- Scene ids are for your metadata. Duplicates are your problem, so generate them from your own step table.
- Read the returned start and end times rather than assuming the numbers you planned; they are the source for your chapters.
- Later scene stills are pose-anchored continuations of the first frame when scenes share one background, and a request can resolve to only one avatar and one shared scene.
- Silence scenes (
voice.type: "silence") need adurationand make good pauses for a viewer to act, such as during a click-through.
When separate clips are better
If steps change independently, render each as its own 4-60 second clip. Then fixing one step is one re-render instead of redoing a whole video, because a changed script needs a new preview anyway. Use the single-clip chapter approach for short, stable sequences, and see the previews page for how stills are produced.
Sources
Related posts
More in Sume Avatar 1.0
- One voice across 23 languages: MAI-Voice-2.1 vs a Sume avatar voice
MAI-Voice-2.1 keeps one voice across 23 languages. A Sume voice has one primary language and a 409 guard on mismatch. What that means for avatars.
- Pick a stock avatar by avoid_for and brand_safety_notes, not looks
Sume's avatar catalog returns profile metadata with best_for, avoid_for, brand_safety_notes and casting_notes. Read them before you cast a presenter for a clip.
- Griffin-Lite 26 of 54 Turing result: what the sample size says
Tavus reports 26 of 54 callers fooled by Griffin-Lite after a one-minute call. A 95% interval is about 35% to 61%. Python computes it, and says what to claim.
- Tavus Video to Face replica vs a Sume avatar from a photo or prompt
Tavus builds a replica from video or a photo; Sume builds an avatar from a prompt, props or a public photo URL. What each input gives you, and what it costs.
Written by Sume