TTS output_format: choose wav or mp3 for joins, captions and clips
Sume TTS returns mp3 by default. Choose wav when you will join, slice per sentence or feed lip-sync; mp3 is smaller but adds padding at every edge.

Pick wav when the TTS audio will be joined, sliced or used to drive a lip-sync clip, and keep the default mp3 when the file goes straight to a listener. Sume TTS defaults to mp3 at 44,100 Hz and 128 kbps; wav and raw containers are available, with sample rates from 8,000 to 48,000 Hz.
The reasons are in two places: the API reference says per-segment audio files are produced only for wav or raw, and the Timeline audio docs say mp3 adds priming padding at every edge, so wav is the format to keep if you will join the file again.
What are the choices?
output_format sets the container, the encoding such as pcm_s16le, and the sample rate. If you omit it you get mp3. The values below are the ones the schema and docs state.
| Need | Format | Why |
|---|---|---|
| Final file for a listener | mp3 (default) | Smaller; 44.1 kHz, 128 kbps |
| Concat or split on a timeline | wav | Sample-exact; no extra padding |
| Per-sentence audio files | wav or raw | emit_audio clips exist only for these |
| Feeding a lip-sync clip | wav | Clean edges; check the route's size limits |
| Captions timings only | Either | Word timings come from timestamps.words |
Sentence segments
With timestamps.words: true and segmentation: {mode: "sentence"}, the result carries gapless segments. Set emit_audio on a wav or raw job and each segment gets its own audio_url. On mp3 you get the timings but no per-segment files. boundary_lead_ms defaults to 70 and ranges from 0 to 500; it sets how far a segment runs past its last word.
- Need individual lines (game barks, IVR prompts): wav with
emit_audio. - Need only subtitles: mp3 plus word timings.
- Need both: wav, then convert the final mix yourself.
Does format change the price?
Not in the schema read here: the TTS price is $0.0475 per 1,000 characters, rounded up to a cent per job, and the format fields are not listed as price inputs. A larger wav does cost more to store and download, so keep working files only as long as you need them.
| Characters | Math | Charged |
|---|---|---|
| 600 | 0.6 x $0.0475 = $0.0285 | $0.03 |
| 2,000 | 2 x $0.0475 = $0.095 | $0.10 |
| 20,000 | 20 x $0.0475 = $0.95 | $0.95 |
A simple rule
Generate in wav while you are still editing, and encode to mp3 once at the end. If you generate mp3 first and later need to join parts, you will either accept padding at each seam or pay for a second render.
Common mistakes
The first mistake is asking for mp3 and then expecting per-sentence files. Segment audio is produced only for wav or raw, so an mp3 job returns timings and nothing more; switch the format and rerun. The second is changing the sample rate between parts that you plan to concatenate: keep one rate and container for all parts of a narration so the join does not have to resample.
The third is forgetting language on non-English scripts, which is independent of format but shows up at the same time, when you listen to the first render. Check the language, then the format, then the voice.
- Same container, rate and voice for every part of a joined narration.
- Re-run with wav, not mp3, when you need per-segment files.
- Keep the first render short while you pick settings.
Sources
Related posts
More in Developers
- Twenty image jobs, one webhook: mode webhook plus a Python verifier
Submit 20 image requests with mode webhook, receive signed job.completed callbacks, and verify them in Python. Retries, replay window and the poll fallback.
- Unit-test a Sume job poll loop with a fake clock and no network
Inject the status reader and the sleep function to test a Sume poll loop in milliseconds: next_poll_after_seconds, backoff fallback and the client deadline.
- Upgraded your Sume plan but ratelimit-limit is still the old number?
A plan change can take up to 60 seconds to reach the per-key rate limit, because the tier is cached. Why ratelimit-limit lags, and what changes at once.
- uv run a single-file Python script against the Sume API (PEP 723)
A one-file Sume script with inline PEP 723 dependencies runs with uv run and no virtualenv. Submit, poll next_poll_after_seconds, print the result.
Written by Sume