Timeline audio concat or audio.parts: which join for narration takes?

Use audio.parts when the narration only feeds one render; use Timeline audio concat when you need a reusable file. Concat is $0.01 per job, up to 1,800 s.

5 min readSume
All posts

Pick by what you need afterwards

If the joined narration is needed only inside one video render, pass the takes as audio.parts[] in Timeline 1.0. If you need a standalone, reusable audio file with a URL, run Timeline audio with operation: concat. The two joins are both sample-domain with no re-TTS.

The Sume docs state the rule directly: sliced VO inside one render uses audio.parts; a reusable merged file uses Timeline audio.

Side by side

Two ways to join takes (from Sume docs, read 2026-10-08)
Propertyaudio.parts[] in Timeline 1.0Timeline audio concat
ResultAudio inside one rendered MP4A reusable media.sume.com audio file
Limit on pieces20 slicesNot stated on the page read
Output length1 to 1,800 s (audio.duration_seconds)Up to 1,800 s
PriceIncluded in the render: $0.10 per ceil(output minute)$0.01 flat per job
PollingJob envelopeJob envelope: GET /v1/jobs/:id/status and /result

Worked example: six takes of a 12-minute narration

Say six 2-minute takes feed a 12-minute video. With audio.parts the narration join is free inside the render, and the render costs ceil(12) x $0.10 = $1.20. With a standalone concat first you add $0.01 and then render with audio.url, for $1.21 total, in exchange for a file you can reuse for captions, a podcast feed or a second cut.

The extra cent buys the artifact; decide whether you will reuse it.

Failure modes to know

Concat needs parts and rejects ranges or url; split needs url and ranges. Concat parts must share a channel layout, or the worker returns audio_parts_channel_mismatch. Off-host URLs are rejected, so import files first. Refusal codes are listed in the Timeline audio page; the render-side rules are in Timeline 1.0.

A quick decision rule

Ask whether the audio will be used anywhere except this one render. If yes, concat first. If you are not sure, concat anyway: one cent is cheap insurance against regenerating or re-joining later, and the durable file can be used for captions, a second video cut or an audio-only release.

If the narration is more than 20 slices, audio.parts will not take it, so you would concat in groups first and then pass the groups on.

Gain and trimming

Timeline 1.0 lets you set audio.gain_db between minus 60 and 12 for a single spine, and source_in lets you start partway into a single file. Those do not combine with parts, per the docs, so if you need to trim the start of a joined narration, trim before joining or concat first and then use source_in on the resulting single file.

Sources

More in Developers

All Developers posts

Written by Sume