26 dialogue lines, a 20-part concat cap: two $0.01 jobs, then offsets

Timeline audio concat takes 1-20 parts. A 26-line dialogue needs two concat jobs, $0.02 total, plus about $0.17 of TTS. How to chain them and keep the offsets.

5 min readSume
All posts

Sume's timeline audio concat accepts 1 to 20 parts per job, so a 26-line dialogue takes two jobs: lines 1-20 in the first, then the first result plus lines 21-26 (seven parts) in the second. The two concat jobs cost $0.02. If each line is 140 characters, the 26 TTS jobs add 3,640 x $0.0475 / 1,000 = $0.173, for about $0.19 in all.

The chain

Concat joins in the sample domain, with no re-synthesis and no silence at the seams. Every URL must be this workspace's media.sume.com audio, and all parts must share a channel layout or the job fails with audio_parts_channel_mismatch. The output of a concat job is itself a media.sume.com file, so it can be a part of the next one.

Two-job concat for 26 lines, rates read 2026-10-09
JobPartsResultPrice
Concat Alines 1-20 (20 parts)one file, segments[] for lines 1-20$0.01
Concat Bfile A + lines 21-26 (7 parts)one file, segments[] for A and lines 21-26$0.01
TTS26 jobs x 140 characters3,640 characters$0.173
Total$0.193

Keeping the start times

Each concat result returns segments[] with index, start and duration_seconds. Those are the offsets you use to re-base Timeline video[].start, per the docs. After job B, segment 0 is the whole of file A, so the lines inside A keep the starts that job A gave you, and lines 21-26 take their starts from job B, which already include A's full length.

Write the two segment tables into your own record at once. The job result is the only place those numbers exist.

A worked example of the offsets

Suppose lines 1-20 add up to 61.4 seconds (an assumed figure for illustration). Job A returns twenty segments that start at 0 and end at 61.4. Job B takes file A as part 0 and lines 21-26 as parts 1-6. Its segments report part 0 as start 0 with duration 61.4, then line 21 at start 61.4, and so on. The final file is one gapless spine, and every Timeline video[].start comes straight from these numbers.

Plan the render before you pay for it. The unbilled POST /v1/timeline-1.0/plan call returns billable_minutes; a 90-second spine reserves 2 minutes, which is $0.20 at $0.10 per output minute.

Choices that avoid a second job

There are four ways to stay under the cap or make the second job unnecessary.

  • Cap the dialogue at 20 lines per scene. A scene break is a natural place for a cut anyway.
  • If you only need the lines inside one render, use audio.parts[] on the Timeline render, which also takes up to 20 parts, so it has the same ceiling.
  • Use wav for any file you will join again. The docs say mp3 adds priming padding at every edge, which is what makes joins audible.
  • Keep the same output format and sample rate for every TTS line so the channel layouts match.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume