Map avatar video scene ids to onboarding steps with scene_previews

Give each video_inputs scene a stable id and Sume returns it in scene_previews with start, end and duration, so one clip can drive per-step chapters.

4 min readSume
All posts

Set video_inputs[].id to your own stable step key, such as step-1-connect-account, and Sume echoes it back in scene_previews[].id along with start_time_seconds, end_time_seconds, and duration_seconds. That gives you chapter markers and per-step stills for one onboarding clip without keeping your own bookkeeping of which sentence landed where.

What the API accepts and returns

In a multi-scene request, each video_inputs item takes an optional id (1-100 characters, described as a client-provided stable scene id for metadata) and a required voice. The array holds 1-20 scenes, and each spoken scene has up to 2,000 characters of script or input_text. The total planned duration must still fall in the 4-60 second window. A second per-scene id use, scene_previews[].id, is documented as the optional caller-provided id when available from video_inputs[].

Scene fields for step mapping (Sume OpenAPI, read 2026-10-05)
FieldWhereUse
video_inputs[].idRequestYour step key, 1-100 characters
scene_previews[].indexResponsePosition of the scene
scene_previews[].idResponseYour id when available
scene_previews[].start_time_secondsResponseChapter start
scene_previews[].duration_secondsResponseScene length
scene_previews[].preview_image_urlResponseFirst-frame still for the scene

Build chapters from the response

Rather than a separate clip for each step, you can render one short onboarding video with several steps and still link into it by time. This works only up to the 60-second ceiling, so a long checklist should be split by topic into several jobs; that limit and the split strategy are covered by the one-job-per-step approach in the related posts.

Keep a lookup of your step keys. When the resource is ready, read scene_previews, drop entries without an id, and write a chapter list for your player.

def chapters(scene_previews):
    out = []
    for s in scene_previews or []:
        if not s.get("id"):
            continue
        out.append({
            "step": s["id"],
            "start": s.get("start_time_seconds"),
            "still": s.get("preview_image_url"),
        })
    return out

sample = [
    {"index": 0, "id": "step-1", "start_time_seconds": 0,
     "preview_image_url": "https://media.sume.com/a.png"},
    {"index": 1, "id": "step-2", "start_time_seconds": 9.5,
     "preview_image_url": None},
]
print(chapters(sample))

Checks before you rely on it

  • Scene ids are for your metadata. Duplicates are your problem, so generate them from your own step table.
  • Read the returned start and end times rather than assuming the numbers you planned; they are the source for your chapters.
  • Later scene stills are pose-anchored continuations of the first frame when scenes share one background, and a request can resolve to only one avatar and one shared scene.
  • Silence scenes (voice.type: "silence") need a duration and make good pauses for a viewer to act, such as during a click-through.

When separate clips are better

If steps change independently, render each as its own 4-60 second clip. Then fixing one step is one re-render instead of redoing a whole video, because a changed script needs a new preview anyway. Use the single-clip chapter approach for short, stable sequences, and see the previews page for how stills are produced.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume