Keep the sound of AI clips in a Sume timeline: detach to a spine
A Timeline 1.0 render takes one audio spine. To keep audio generated with your clips, detach each clip's track, join the parts, and use that as the spine.

If your AI clips came back with generated sound, the documented way to keep it in a stitched film on Sume is to detach each clip's audio, join the pieces, and pass the result as the timeline's audio spine. Timeline 1.0 is defined around one spine (or silence), and the docs do not describe mixing the clips' own audio into the render.
I could not find documentation that says what happens to a video slot's embedded audio, so this post does not claim either way. It gives you the route the docs do describe, which is deterministic.
What does Timeline 1.0 do with audio?
A render requires audio.duration_seconds plus either audio.url, audio.parts[], or audio.mode: "silence". The result is one MP4 whose length matches the spine. An optional soundtrack adds a music bed with gain_db, loop, fade_out_seconds and duck_db, but ducking needs a real spine, not silence.
Generation is where the sound comes from. On POST /v1/videos, generate_audio defaults to the model's audio capability, so a model with native audio returns clips that carry a track unless you turn it off.
How do I build a spine from the clips' own audio?
Step one is audio_detach per clip: $0.01 per job, format wav by default (sample-exact), channels source by default, and the output keeps the source rate and channels unless you set them. Keep wav and source channels, because the render keeps whatever fidelity the spine has. Sume even warns with audio_spine_low_fidelity when the spine is low-rate or mono and points you to audio_detach output at source rate.
Step two is the join. Put the detached files in audio.parts[] of the render itself (up to 20 gapless slices, each url with optional source_in and duration, sample-domain join, no re-synthesis). If you want a reusable file, run POST /v1/timeline-1.0/audio with operation: "concat" ($0.01 flat) and use the returned audio_url as audio.url. Parts must share one channel layout, or concat fails with audio_parts_channel_mismatch.
{
"audio": {
"duration_seconds": 15,
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/s1.wav", "duration": 5 },
{ "url": "https://media.sume.com/artifacts/artf_demo/s2.wav", "duration": 5 },
{ "url": "https://media.sume.com/artifacts/artf_demo/s3.wav", "duration": 5 }
]
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/s1.mp4", "start": 0, "duration": 5 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/s2.mp4", "start": 5, "duration": 5 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/s3.mp4", "start": 10, "duration": 5 }
]
}What does this cost and what limits apply?
Sum for eight shots under a minute: $0.08 for detaching plus $0.10 for the render, $0.18 before any generation cost. Confirm the live rates in GET /v1/catalog.
| Step | Rate | Limit |
|---|---|---|
| audio_detach, 8 clips | $0.01 per job, $0.08 total | Output up to 900 s per job |
| audio.parts[] in the render | No extra job | Up to 20 slices, gapless |
| Optional timeline audio concat | $0.01 flat per job | Parts 1 to 20, output up to 1800 s |
| Timeline render, under 1 minute | $0.10 per output minute, rounded up | Spine 1 to 1800 s, 1 to 200 slots |
What can break the sync?
Each detached part must be as long as its slot. Set the slot start values as the running total of the part durations, and check that the parts sum to at least audio.duration_seconds, or the render is refused with audio_parts_shorter_than_duration. Probe each clip with video inspect first if you are not sure the generated duration is exactly what you requested.
Crossfades are compensated by the compiler, so the declared starts stay authoritative and the audio does not shift. A clip with no audio track makes audio_detach fail; check for audio before you detach, using the probe.
Finally, a caution about dialogue: Sume's own agent guidance says video models do not lip-sync, so if a shot needs a person speaking exact words, make the voice separately (TTS, then a talking-clip job) and use it as the spine instead of detaching a generated track.
Sources
Related posts
More in Media tools
- Reference ingest OCR needs_verification: read the crop
Low-confidence on-screen text from reference ingest returns as needs_verification with a native crop. How the 0.85 default works and how to read the manifest.
- reference_ingest_semantic_unavailable: what to do instead
semantic: true is refused on reference ingest, and reference_ingest_unavailable means the media runtime lacks the function. The two errors and what to call.
- Restyle burned-in captions without paying for a second transcription
Pass source_caption_id to POST /v1/video-captions to re-burn the same video in another style, reusing its word timings. Billing is still one render.
- Revert an AI video edit: why Sume trims never touch your source
Descript moved its Revert button next to the AI response. In an API pipeline revert is free: each Sume edit returns a new artifact and the source stays put.
Written by Sume