26 dialogue lines, a 20-part concat cap: two $0.01 jobs, then offsets
Timeline audio concat takes 1-20 parts. A 26-line dialogue needs two concat jobs, $0.02 total, plus about $0.17 of TTS. How to chain them and keep the offsets.

Sume's timeline audio concat accepts 1 to 20 parts per job, so a 26-line dialogue takes two jobs: lines 1-20 in the first, then the first result plus lines 21-26 (seven parts) in the second. The two concat jobs cost $0.02. If each line is 140 characters, the 26 TTS jobs add 3,640 x $0.0475 / 1,000 = $0.173, for about $0.19 in all.
The chain
Concat joins in the sample domain, with no re-synthesis and no silence at the seams. Every URL must be this workspace's media.sume.com audio, and all parts must share a channel layout or the job fails with audio_parts_channel_mismatch. The output of a concat job is itself a media.sume.com file, so it can be a part of the next one.
| Job | Parts | Result | Price |
|---|---|---|---|
| Concat A | lines 1-20 (20 parts) | one file, segments[] for lines 1-20 | $0.01 |
| Concat B | file A + lines 21-26 (7 parts) | one file, segments[] for A and lines 21-26 | $0.01 |
| TTS | 26 jobs x 140 characters | 3,640 characters | $0.173 |
| Total | $0.193 |
Keeping the start times
Each concat result returns segments[] with index, start and duration_seconds. Those are the offsets you use to re-base Timeline video[].start, per the docs. After job B, segment 0 is the whole of file A, so the lines inside A keep the starts that job A gave you, and lines 21-26 take their starts from job B, which already include A's full length.
Write the two segment tables into your own record at once. The job result is the only place those numbers exist.
A worked example of the offsets
Suppose lines 1-20 add up to 61.4 seconds (an assumed figure for illustration). Job A returns twenty segments that start at 0 and end at 61.4. Job B takes file A as part 0 and lines 21-26 as parts 1-6. Its segments report part 0 as start 0 with duration 61.4, then line 21 at start 61.4, and so on. The final file is one gapless spine, and every Timeline video[].start comes straight from these numbers.
Plan the render before you pay for it. The unbilled POST /v1/timeline-1.0/plan call returns billable_minutes; a 90-second spine reserves 2 minutes, which is $0.20 at $0.10 per output minute.
Choices that avoid a second job
There are four ways to stay under the cap or make the second job unnecessary.
- Cap the dialogue at 20 lines per scene. A scene break is a natural place for a cut anyway.
- If you only need the lines inside one render, use
audio.parts[]on the Timeline render, which also takes up to 20 parts, so it has the same ceiling. - Use wav for any file you will join again. The docs say mp3 adds priming padding at every edge, which is what makes joins audible.
- Keep the same output format and sample rate for every TTS line so the channel layouts match.
Sources
Related posts
More in Developers
- A 28-minute talk transcribed on Sume: 3 detach ranges plus STT = $0.31
Audio detach caps output at 900 s and STT estimates cap at 10 minutes. Here is how a 28-minute talk splits into three ranges, and what it costs: $0.31.
- 34 pause cuts, 20 audio parts: split the spine in two levels
Timeline audio.parts holds 20 slices. For 34 pause cuts, build two concat files of 17, feed them as two parts, and re-base each video slot start.
- 37 video jobs submitted at once: which Sume plan accepts them all
Free accepts 6 of 37 video jobs, Pro 24, Startup all 37 with 11 spare slots, Scale all 37. Concurrency, queue and accepted-capacity table from the Sume docs.
- 4 of Sume's 17 aspect ratios exceed 3:1, GPT Image 2.5's size cap
Sume normalizes 17 aspect ratios; 1:4, 4:1, 1:8 and 8:1 are wider than the 3:1 limit on GPT Image 2.5 custom sizes. A check script and the 2 other rules.
Written by Sume