TTS output_format: choose wav or mp3 for joins, captions and clips

Sume TTS returns mp3 by default. Choose wav when you will join, slice per sentence or feed lip-sync; mp3 is smaller but adds padding at every edge.

5 min readSume
All posts

Pick wav when the TTS audio will be joined, sliced or used to drive a lip-sync clip, and keep the default mp3 when the file goes straight to a listener. Sume TTS defaults to mp3 at 44,100 Hz and 128 kbps; wav and raw containers are available, with sample rates from 8,000 to 48,000 Hz.

The reasons are in two places: the API reference says per-segment audio files are produced only for wav or raw, and the Timeline audio docs say mp3 adds priming padding at every edge, so wav is the format to keep if you will join the file again.

What are the choices?

output_format sets the container, the encoding such as pcm_s16le, and the sample rate. If you omit it you get mp3. The values below are the ones the schema and docs state.

Format choice (Sume schema and docs, checked 2026-10-10)
NeedFormatWhy
Final file for a listenermp3 (default)Smaller; 44.1 kHz, 128 kbps
Concat or split on a timelinewavSample-exact; no extra padding
Per-sentence audio fileswav or rawemit_audio clips exist only for these
Feeding a lip-sync clipwavClean edges; check the route's size limits
Captions timings onlyEitherWord timings come from timestamps.words

Sentence segments

With timestamps.words: true and segmentation: {mode: "sentence"}, the result carries gapless segments. Set emit_audio on a wav or raw job and each segment gets its own audio_url. On mp3 you get the timings but no per-segment files. boundary_lead_ms defaults to 70 and ranges from 0 to 500; it sets how far a segment runs past its last word.

  • Need individual lines (game barks, IVR prompts): wav with emit_audio.
  • Need only subtitles: mp3 plus word timings.
  • Need both: wav, then convert the final mix yourself.

Does format change the price?

Not in the schema read here: the TTS price is $0.0475 per 1,000 characters, rounded up to a cent per job, and the format fields are not listed as price inputs. A larger wav does cost more to store and download, so keep working files only as long as you need them.

Cost for common script sizes (rate checked 2026-10-10)
CharactersMathCharged
6000.6 x $0.0475 = $0.0285$0.03
2,0002 x $0.0475 = $0.095$0.10
20,00020 x $0.0475 = $0.95$0.95

A simple rule

Generate in wav while you are still editing, and encode to mp3 once at the end. If you generate mp3 first and later need to join parts, you will either accept padding at each seam or pay for a second render.

Common mistakes

The first mistake is asking for mp3 and then expecting per-sentence files. Segment audio is produced only for wav or raw, so an mp3 job returns timings and nothing more; switch the format and rerun. The second is changing the sample rate between parts that you plan to concatenate: keep one rate and container for all parts of a narration so the join does not have to resample.

The third is forgetting language on non-English scripts, which is independent of format but shows up at the same time, when you listen to the first render. Check the language, then the format, then the voice.

  • Same container, rate and voice for every part of a joined narration.
  • Re-run with wav, not mp3, when you need per-segment files.
  • Keep the first render short while you pick settings.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume