Timeline audio concat or audio.parts: when the $0.01 job is wasted

Joining audio inside one render is free with audio.parts[] (up to 20 slices). The $0.01 concat job pays off only for a reusable file. A decision table.

4 min readSume
All posts

If you need to join audio for one Timeline render, use audio.parts[] inside the render request, which adds no job and no charge beyond the $0.10 per started minute you already pay. Run the $0.01 Timeline audio concat job only when you need the joined audio as a reusable file, for example as an Avatar 1.0 audio input or for several renders.

Both routes are described on the Sume docs pages, and the Timeline audio page itself says: if you need a join only inside one render, use audio.parts[] and do not use the concat job. This post is the decision table, with the cent counted.

What each route gives you

Both do a sample-domain join with no silence at the seams and no re-synthesis. The differences are in what you get back and what you pay.

Inline audio.parts[] against the concat job, Sume docs read 2026-10-08
Propertyaudio.parts[] in a renderConcat job
PriceIncluded in the render, $0.10 per started output minute$0.01 per job
Parts allowedUp to 20 slices1 to 20 parts
OutputOnly the rendered MP4A durable audio file with audio_url, duration_seconds, segments[]
Reusable in other jobsNoYes (render, Avatar 1.0 audio, other renders)
Offsets for video startsYou compute themReturned as segments[] with start and duration_seconds
FormatNot a separate filewav (default) or mp3

The cent that matters

The concat job costs one cent, so the money is not the issue. The issue is extra steps: a job to poll, a file to store, and an Idempotency-Key to manage. For a one-off render, the inline route has fewer moving parts. For a campaign of 24 renders that share the same 6-part voice and bed join, the concat route costs $0.01 once instead of describing 6 parts 24 times.

The returned segments[] are the useful part. They give the concat offsets you use to re-base video[].start in the render, so each clip begins where its audio does. With inline parts you do that arithmetic yourself.

A worked case

Take a 40-second Short with a 12-second intro bed, a 20-second main bed and an 8-second outro bed, all Lyria files you already hold. Inline, you list three parts in audio.parts[] with duration 12, 20 and 8, and the render is $0.10. With the concat job, you pay $0.01, wait for the job, then pass the returned audio_url as audio.url, for a total of $0.11 and one more call. The extra cent buys the file, so it is only worth it if a second render, an Avatar 1.0 clip or another market version will use the same join.

Another reason to prefer the job is auditing. Its result lists segments[] with the offset of each part, so a reviewer can see where each bed starts without opening the render.

Two limits to remember

The inline route requires that the declared part lengths cover the render: the Timeline page lists audio_parts_shorter_than_duration as the error when the sum is less than audio.duration_seconds. The concat route requires that every part share a channel layout (audio_parts_channel_mismatch), so check the layouts of mixed inputs before you submit.

A last case: a Lyria bed and a spoken line. Join them with neither route if you want the bed under the voice. Use the render's soundtrack field with duck_db (0 to 20), which is made for that, and keep the voice as the spine.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume