Timeline audio concat or audio.parts: which join for narration takes?
Use audio.parts when the narration only feeds one render; use Timeline audio concat when you need a reusable file. Concat is $0.01 per job, up to 1,800 s.

Pick by what you need afterwards
If the joined narration is needed only inside one video render, pass the takes as audio.parts[] in Timeline 1.0. If you need a standalone, reusable audio file with a URL, run Timeline audio with operation: concat. The two joins are both sample-domain with no re-TTS.
The Sume docs state the rule directly: sliced VO inside one render uses audio.parts; a reusable merged file uses Timeline audio.
Side by side
| Property | audio.parts[] in Timeline 1.0 | Timeline audio concat |
|---|---|---|
| Result | Audio inside one rendered MP4 | A reusable media.sume.com audio file |
| Limit on pieces | 20 slices | Not stated on the page read |
| Output length | 1 to 1,800 s (audio.duration_seconds) | Up to 1,800 s |
| Price | Included in the render: $0.10 per ceil(output minute) | $0.01 flat per job |
| Polling | Job envelope | Job envelope: GET /v1/jobs/:id/status and /result |
Worked example: six takes of a 12-minute narration
Say six 2-minute takes feed a 12-minute video. With audio.parts the narration join is free inside the render, and the render costs ceil(12) x $0.10 = $1.20. With a standalone concat first you add $0.01 and then render with audio.url, for $1.21 total, in exchange for a file you can reuse for captions, a podcast feed or a second cut.
The extra cent buys the artifact; decide whether you will reuse it.
Failure modes to know
Concat needs parts and rejects ranges or url; split needs url and ranges. Concat parts must share a channel layout, or the worker returns audio_parts_channel_mismatch. Off-host URLs are rejected, so import files first. Refusal codes are listed in the Timeline audio page; the render-side rules are in Timeline 1.0.
A quick decision rule
Ask whether the audio will be used anywhere except this one render. If yes, concat first. If you are not sure, concat anyway: one cent is cheap insurance against regenerating or re-joining later, and the durable file can be used for captions, a second video cut or an audio-only release.
If the narration is more than 20 slices, audio.parts will not take it, so you would concat in groups first and then pass the groups on.
Gain and trimming
Timeline 1.0 lets you set audio.gain_db between minus 60 and 12 for a single spine, and source_in lets you start partway into a single file. Those do not combine with parts, per the docs, so if you need to trim the start of a joined narration, trim before joining or concat first and then use source_in on the resulting single file.
Sources
More in Developers
- Timeline plan call: price a narrated render before you spend
POST /v1/timeline-1.0/plan is unbilled and returns billable_minutes and estimated_cost_usd_micros. How to price a narration plus music render first.
- Timeline plan needs no Idempotency-Key; render does: a wrapper
POST /v1/timeline-1.0/plan is free and takes no Idempotency-Key. Render is paid and requires one. A wrapper plans first, then renders with a derived key.
- Transcribe a 90-second clip with video inspect: cost and reservation
Video inspect with transcribe true adds Sume STT at $0.01 per audio minute; send duration_seconds 90 so the hold fits, or it reserves one minute.
- Transcript from a video on Sume: inspect transcribe or detach first?
Sume can transcribe a video through video inspect (1,800 s limit, compute plus $0.01 per minute) or from a detached 16 kHz mono wav. How to choose.
Written by Sume