TTS per sentence, then concat: join up to 20 clips in one file

Generate TTS one sentence at a time, then join up to 20 clips with Sume timeline audio concat: sample-domain, no re-synthesis, $0.01 flat per job.

4 min readSume
All posts

To build one long voice file from per-sentence TTS clips, import the clips and call timeline audio with operation: "concat" and up to 20 ordered parts[]. Sume's docs describe the join as sample-domain, with no re-TTS and no silence at the seams, at a public rate of $0.01 flat per job.

Why split at all? The Deepgram changelog describes Flux TTS inline pauses of 500 to 3000 ms, up to 8 per request, so long scripts end up spread across requests. Sume facts are from Timeline audio, read 2026-10-01.

Why would I generate TTS one sentence at a time?

Short requests are easy to retry and to re-take. A single TTS request on Sume accepts up to 20000 characters, so per-sentence generation is a choice, not a limit. When one sentence needs a new take, you regenerate that clip and rejoin.

What does concat accept?

Concat rules from the Sume docs read 2026-10-01.
RuleDetail
Partsparts[], 1 to 20, ordered; each { url, source_in?, duration? }
SourceAlready this workspace's media.sume.com audio; import first
ChannelsParts must share one channel layout (audio_parts_channel_mismatch)
Max output1800 seconds
ResultOne audio_url plus segments[] with start offsets

What if I have more than 20 sentences?

Join in batches, then join the batch outputs. Keep wav, the default, for every intermediate file: mp3 re-adds priming padding at each edge. Use segments[] from the result to re-base any video that follows the audio.

Is a separate join job needed at all?

Not if the join only matters inside one render: the docs say to use Timeline 1.0 audio.parts[] and skip this job. Use concat when you want a reusable file. See also sentence segments in TTS.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume