Use timeline audio concat segments to set Timeline video starts

After a Sume timeline audio concat, copy each segments[].start into video[].start. A Python snippet, plus the rules video[0].start must be 0.

6 min readSume
All posts

How do you line up video cuts with a concatenated voiceover? Read segments[] from the timeline audio result and use each segment's start as the matching video[].start in the Timeline 1.0 render. The Sume timeline audio docs describe the segments as the concat offsets to re-base video[].start against.

Each segment has an index, a start and a duration_seconds, so one voiceover line maps to one clip.

The mapping

The concat joins up to 20 parts in the sample domain, with no re-synthesis and no silence at the seams, and returns one audio_url. The Timeline render then takes that file as audio.url and an array of clips. Two rules from the Timeline 1.0 docs matter here.

Fields that connect timeline audio to Timeline 1.0 (read 2026-10-03)
Source fieldTarget fieldRule
segments[i].startvideo[i].startvideo[0].start must be 0; later starts must increase
segments[i].duration_secondsvideo[i].durationOn-screen length is at least 0.2 s
audio_urlaudio.urlOne Sume-hosted spine, exclusive with audio.parts
duration_secondsaudio.duration_secondsRequired, from 1 to 1800 seconds

A short Python builder

This builds the render body from a concat result. The result literal below stands in for a real job result; the file URLs are placeholders for your own media.sume.com artifacts.

result = {
    "audio_url": "https://media.sume.com/artifacts/artf_demo/vo.wav",
    "duration_seconds": 14.0,
    "segments": [
        {"index": 0, "start": 0.0, "duration_seconds": 4.5},
        {"index": 1, "start": 4.5, "duration_seconds": 5.0},
        {"index": 2, "start": 9.5, "duration_seconds": 4.5},
    ],
}
clips = [
    "https://media.sume.com/artifacts/artf_demo/a.mp4",
    "https://media.sume.com/artifacts/artf_demo/b.mp4",
    "https://media.sume.com/artifacts/artf_demo/c.mp4",
]
assert len(clips) == len(result["segments"])
assert result["segments"][0]["start"] == 0
body = {
    "audio": {
        "url": result["audio_url"],
        "duration_seconds": result["duration_seconds"],
    },
    "video": [
        {"source_url": u, "start": s["start"],
         "duration": s["duration_seconds"]}
        for u, s in zip(clips, result["segments"])
    ],
}
print(body["video"][1])

Why not compute starts yourself

You could add up the durations of the parts you submitted, but the segments array is what the server measured after the join. Source files that are slightly longer or shorter than their declared duration will drift if you sum your own numbers. Use the returned offsets.

The Timeline docs say declared starts are authoritative and that coverage may trail the spine by at most 0.5 seconds. So the last clip should end at or very near the audio length.

When to skip the separate job

A concat job mints a reusable file for a flat $0.01. If the join is only needed inside one render, the Timeline docs say to put the slices on audio.parts[] and skip the job. Then you do not get a segments array back, so compute start from the part durations you declared, and use the separate job when you need the measured offsets or want to reuse the merged audio elsewhere.

Timeline rendering is billed at $0.10 per ceil of output minute, so a 14 second video is one billable minute.

Common mistakes

Parts must share one channel layout, or the job fails with audio_parts_channel_mismatch; the channel mismatch fix covers it. Parts and ranges must already be this workspace's media.sume.com audio, so import first. Both requests need an Idempotency-Key. Poll the job envelope, since the concat has no GET of its own.

One last check before you render: add up the segment durations and compare the total with the audio duration. They should agree to within a frame or two. If they do not, a source part was probably trimmed by source_in or duration and you should re-read the offsets rather than patch them by hand.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume