One voice-over under ten AI shots: Timeline spine and start times

Put one narration track under ten AI video shots on Sume Timeline: video[0].start is 0, later starts increase, and the last shot may end up to 0.5 s early.

5 min readSume
All posts

To lay one voice-over under ten AI shots on Sume Timeline, send one audio.url with its duration_seconds, then a video[] list in which the first slot starts at 0 and every later start is larger than the one before. Each slot has a duration, and the sum of the slots should cover the narration: coverage can stop at most 0.5 seconds before the end of the spine. A 60 second track is one billable minute, $0.10.

This is the shape behind long multi-shot work such as the 10-keyframe plans in the fal explainer for Kling 4.0 (read 2026-10-05). On Sume you build it from separate clips and put the narration on the spine.

The rules

The Timeline docs list the rules. audio.duration_seconds is 1 to 1,800 and is the output length. video[0].start must be 0, and later starts must increase. Declared starts are authoritative: the compiler compensates for a transition overlap and does not pre-shift your starts. Each duration is at least 0.2 seconds.

Every URL must be an artifact or asset in your workspace on media.sume.com, so import any outside file first with POST /v1/media-imports. A render call needs an Idempotency-Key. A render defaults to async, and mode: sync waits up to 30 seconds.

Two more rules are worth knowing. A plan call is unbilled and needs no Idempotency-Key, but it cannot predict the short-source pad and loop warnings, which only a real render reports. A soundtrack bed is a separate field from the spine, so a music track can sit under the voice-over with its own gain, and the voice-over stays the clock that sets the output length.

Ten shots over a 60 second narration

Timeline video[] slots for ten shots over a 60 s spine, read 2026-10-05
SlotstartdurationTransition
106none (first slot)
266optional
3126optional
4-918 to 486 eachoptional
10546optional

Matching shot length to slot length

If the shots come from AI jobs, request each at the length it needs on screen, or longer, and trim on the slot with source_in and duration. A model with a 4 second floor, such as Seedance 2.5 or Kling 3, can fill a 6 second slot only if you request 6 seconds, but a Wan 3.0 clip of 2 seconds is a good fit for a faster cut.

If a clip is shorter than its slot, the job pads or loops it and reports a soft warning. That is not a failure, but it usually looks wrong, so check the warnings array in the result.

Ten shots over 60 seconds is a pace of one cut every six seconds. If your narration is faster, say 90 seconds over ten shots, the slots are nine seconds, and a model with a 10 second ceiling such as Omni can fill each one. If it is slower, use fewer, longer shots. The slot lengths do not have to be equal: sum them and make sure the total reaches the spine.

Build the slots

This builds the slots from a list of shot lengths and checks that they cover the spine. It prints the render body without sending it.

The check in the script is the same one the API makes: if the slots stop more than half a second before the end of the narration, the render is rejected, so catching it in your own code gives you a clearer message earlier. Keep the start times as plain running totals, because declared starts are authoritative and Timeline will not shift them for you.

import json

lengths = [6, 6, 6, 6, 6, 6, 6, 6, 6, 6]
spine = 60
slots, t = [], 0
for i, d in enumerate(lengths, 1):
    slots.append({
        "source_url": f"https://media.sume.com/artifacts/artf_demo/shot{i:02d}.mp4",
        "start": t,
        "duration": d,
    })
    t += d
assert t >= spine - 0.5, "slots stop more than 0.5 s before the spine"
body = {
    "audio": {
        "url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
        "duration_seconds": spine,
    },
    "video": slots,
}
print(len(slots), "slots, covers", t, "s of", spine)
print(json.dumps(body)[:160])

Plan, then render

Run POST /v1/timeline-1.0/plan before the render. It is unbilled, and it returns segment_count, billable_minutes and estimated_cost_usd_micros. Then render, poll the job, and read video_url from GET /v1/jobs/{id}/result. The jobs docs describe the routes. For a canvas, set output.width and output.height yourself, because the default is 1080 x 1920.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume