One voice-over under ten AI shots: Timeline spine and start times
Put one narration track under ten AI video shots on Sume Timeline: video[0].start is 0, later starts increase, and the last shot may end up to 0.5 s early.

To lay one voice-over under ten AI shots on Sume Timeline, send one audio.url with its duration_seconds, then a video[] list in which the first slot starts at 0 and every later start is larger than the one before. Each slot has a duration, and the sum of the slots should cover the narration: coverage can stop at most 0.5 seconds before the end of the spine. A 60 second track is one billable minute, $0.10.
This is the shape behind long multi-shot work such as the 10-keyframe plans in the fal explainer for Kling 4.0 (read 2026-10-05). On Sume you build it from separate clips and put the narration on the spine.
The rules
The Timeline docs list the rules. audio.duration_seconds is 1 to 1,800 and is the output length. video[0].start must be 0, and later starts must increase. Declared starts are authoritative: the compiler compensates for a transition overlap and does not pre-shift your starts. Each duration is at least 0.2 seconds.
Every URL must be an artifact or asset in your workspace on media.sume.com, so import any outside file first with POST /v1/media-imports. A render call needs an Idempotency-Key. A render defaults to async, and mode: sync waits up to 30 seconds.
Two more rules are worth knowing. A plan call is unbilled and needs no Idempotency-Key, but it cannot predict the short-source pad and loop warnings, which only a real render reports. A soundtrack bed is a separate field from the spine, so a music track can sit under the voice-over with its own gain, and the voice-over stays the clock that sets the output length.
Ten shots over a 60 second narration
| Slot | start | duration | Transition |
|---|---|---|---|
| 1 | 0 | 6 | none (first slot) |
| 2 | 6 | 6 | optional |
| 3 | 12 | 6 | optional |
| 4-9 | 18 to 48 | 6 each | optional |
| 10 | 54 | 6 | optional |
Matching shot length to slot length
If the shots come from AI jobs, request each at the length it needs on screen, or longer, and trim on the slot with source_in and duration. A model with a 4 second floor, such as Seedance 2.5 or Kling 3, can fill a 6 second slot only if you request 6 seconds, but a Wan 3.0 clip of 2 seconds is a good fit for a faster cut.
If a clip is shorter than its slot, the job pads or loops it and reports a soft warning. That is not a failure, but it usually looks wrong, so check the warnings array in the result.
Ten shots over 60 seconds is a pace of one cut every six seconds. If your narration is faster, say 90 seconds over ten shots, the slots are nine seconds, and a model with a 10 second ceiling such as Omni can fill each one. If it is slower, use fewer, longer shots. The slot lengths do not have to be equal: sum them and make sure the total reaches the spine.
Build the slots
This builds the slots from a list of shot lengths and checks that they cover the spine. It prints the render body without sending it.
The check in the script is the same one the API makes: if the slots stop more than half a second before the end of the narration, the render is rejected, so catching it in your own code gives you a clearer message earlier. Keep the start times as plain running totals, because declared starts are authoritative and Timeline will not shift them for you.
import json
lengths = [6, 6, 6, 6, 6, 6, 6, 6, 6, 6]
spine = 60
slots, t = [], 0
for i, d in enumerate(lengths, 1):
slots.append({
"source_url": f"https://media.sume.com/artifacts/artf_demo/shot{i:02d}.mp4",
"start": t,
"duration": d,
})
t += d
assert t >= spine - 0.5, "slots stop more than 0.5 s before the spine"
body = {
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
"duration_seconds": spine,
},
"video": slots,
}
print(len(slots), "slots, covers", t, "s of", spine)
print(json.dumps(body)[:160])Plan, then render
Run POST /v1/timeline-1.0/plan before the render. It is unbilled, and it returns segment_count, billable_minutes and estimated_cost_usd_micros. Then render, poll the job, and read video_url from GET /v1/jobs/{id}/result. The jobs docs describe the routes. For a canvas, set output.width and output.height yourself, because the default is 1080 x 1920.
Sources
Related posts
More in Media tools
- Perfume unboxing captions with the scent names spelled right
Fix mis-heard product names in auto-captions by sending script_text: Sume keeps the speech timings and burns your spelling, $0.20 per clip up to 60 seconds.
- Pet adoption video music: a 30-second hopeful shelter bed
Brief a 30-second hopeful acoustic bed for a shelter's pet adoption appeal: one generation and a one-minute render, $0.225 on Sume.
- Pick the cleanest last frame: sample 24 stills before chaining
Before chaining AI clips, sample up to 24 stills with Sume video frames (fps up to 2) and choose the best one as the next first_frame, not just the last.
- Pinterest video ads take H.264 or H.265: do you need H.265?
Pinterest video ads accept H.264 or H.265 in MP4, MOV or M4V, so an H.264 clip is fine. Sume's exact trim uses libx264; confirm any other output with a probe.
Written by Sume