Keep the sound of AI clips in a Sume timeline: detach to a spine

A Timeline 1.0 render takes one audio spine. To keep audio generated with your clips, detach each clip's track, join the parts, and use that as the spine.

5 min readSume
All posts

If your AI clips came back with generated sound, the documented way to keep it in a stitched film on Sume is to detach each clip's audio, join the pieces, and pass the result as the timeline's audio spine. Timeline 1.0 is defined around one spine (or silence), and the docs do not describe mixing the clips' own audio into the render.

I could not find documentation that says what happens to a video slot's embedded audio, so this post does not claim either way. It gives you the route the docs do describe, which is deterministic.

What does Timeline 1.0 do with audio?

A render requires audio.duration_seconds plus either audio.url, audio.parts[], or audio.mode: "silence". The result is one MP4 whose length matches the spine. An optional soundtrack adds a music bed with gain_db, loop, fade_out_seconds and duck_db, but ducking needs a real spine, not silence.

Generation is where the sound comes from. On POST /v1/videos, generate_audio defaults to the model's audio capability, so a model with native audio returns clips that carry a track unless you turn it off.

How do I build a spine from the clips' own audio?

Step one is audio_detach per clip: $0.01 per job, format wav by default (sample-exact), channels source by default, and the output keeps the source rate and channels unless you set them. Keep wav and source channels, because the render keeps whatever fidelity the spine has. Sume even warns with audio_spine_low_fidelity when the spine is low-rate or mono and points you to audio_detach output at source rate.

Step two is the join. Put the detached files in audio.parts[] of the render itself (up to 20 gapless slices, each url with optional source_in and duration, sample-domain join, no re-synthesis). If you want a reusable file, run POST /v1/timeline-1.0/audio with operation: "concat" ($0.01 flat) and use the returned audio_url as audio.url. Parts must share one channel layout, or concat fails with audio_parts_channel_mismatch.

{
  "audio": {
    "duration_seconds": 15,
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/s1.wav", "duration": 5 },
      { "url": "https://media.sume.com/artifacts/artf_demo/s2.wav", "duration": 5 },
      { "url": "https://media.sume.com/artifacts/artf_demo/s3.wav", "duration": 5 }
    ]
  },
  "video": [
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s1.mp4", "start": 0, "duration": 5 },
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s2.mp4", "start": 5, "duration": 5 },
    { "source_url": "https://media.sume.com/artifacts/artf_demo/s3.mp4", "start": 10, "duration": 5 }
  ]
}

What does this cost and what limits apply?

Sum for eight shots under a minute: $0.08 for detaching plus $0.10 for the render, $0.18 before any generation cost. Confirm the live rates in GET /v1/catalog.

Audio spine from clip audio for an eight-shot film (Sume docs, read 2026-10-02)
StepRateLimit
audio_detach, 8 clips$0.01 per job, $0.08 totalOutput up to 900 s per job
audio.parts[] in the renderNo extra jobUp to 20 slices, gapless
Optional timeline audio concat$0.01 flat per jobParts 1 to 20, output up to 1800 s
Timeline render, under 1 minute$0.10 per output minute, rounded upSpine 1 to 1800 s, 1 to 200 slots

What can break the sync?

Each detached part must be as long as its slot. Set the slot start values as the running total of the part durations, and check that the parts sum to at least audio.duration_seconds, or the render is refused with audio_parts_shorter_than_duration. Probe each clip with video inspect first if you are not sure the generated duration is exactly what you requested.

Crossfades are compensated by the compiler, so the declared starts stay authoritative and the audio does not shift. A clip with no audio track makes audio_detach fail; check for audio before you detach, using the probe.

Finally, a caution about dialogue: Sume's own agent guidance says video models do not lip-sync, so if a shot needs a person speaking exact words, make the voice separately (TTS, then a talking-clip job) and use it as the spine instead of detaching a generated track.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume