Join TTS lines into one track: match output_format before concat
Sume timeline audio concat joins 1 to 20 parts but needs one channel layout. Set the same TTS output_format on every line, then concat once.

To join several Sume TTS lines into one voiceover track, send the same output_format on every line, then call timeline audio concat once with the audio urls as parts. Concat takes 1 to 20 parts and all of them must share a channel layout, or the job fails with audio_parts_channel_mismatch.
Match the lines
TTS 1.0 defaults to mp3 at 44100 Hz and 128 kbps. It also supports wav and raw. If every line uses the default, the parts match. Mix a mono and a stereo line and the concat fails. The cleanest rule is to set one output_format in your batch code and reuse it.
| Item | Rule |
|---|---|
| operation | concat |
| parts | 1 to 20 audio urls |
| Top-level url or ranges | Rejected: audio_concat_takes_no_url, audio_concat_takes_no_ranges |
| Channel layout | Must match: audio_parts_channel_mismatch |
| Format | wav or mp3 |
| Price | $0.01 per job |
Longer scripts
Past 20 lines, concat in groups of 20 and concat the group outputs. Or skip the join: Timeline 1.0 takes audio.parts[] for a single render, so you need no separate audio file. Use concat when you need a durable track for another job, such as an avatar video's audio.
Poll the job
Timeline audio has no resource GET. Read GET /v1/jobs/:id/status, then /result. With mode: "sync" you wait at most 30 seconds, otherwise you get a 202.
A note on formats
The audio detach and timeline audio pages describe wav as the sample-exact choice, and mp3 as a smaller one at 128 kbps. When the track will feed another processing step, such as an avatar mux, a lossless line is the safer input. When it only goes into a render, mp3 is smaller to move around.
Whichever you choose, choose once. A mixed batch where some lines came back at a different rate or channel count is the usual reason audio_parts_channel_mismatch appears. The check runs on the worker, so you find out after the job is accepted.
Sources
Related posts
More in Media tools
- TTS mp3 bit rate: 32k to 192k file size per minute for voice audio
Sume TTS 1.0 accepts mp3 bit_rate 32000, 64000, 96000, 128000 or 192000. See the size per minute of each, plus a two-take listening test before you pick one.
- Two-color duotone poster from an AI image with ImageOps.colorize
Turn a Sume image into a duotone: grayscale, autocontrast, then ImageOps.colorize with a dark and a light brand color. Code and ink pairs inside.
- Fades between 2-second AI shots: the 1 s cap and the 8-fade chain
Joining 15 two-second Wan 3.0 shots on Sume Timeline: a fade can be at most 1 s, at most half of the shorter shot, and 8 chained fades are refused. Python.
- Video captions take a public HTTPS URL; other media tools do not
Which URL each Sume video tool accepts: captions fetch a public HTTPS clip, while trim, filter, frames, inspect and timeline need your media.sume.com file.
Written by Sume