Text to speech sample rate: which output_format to pick
Sume's TTS output_format takes mp3, wav or raw, six sample rates from 8000 to 48000 Hz and four PCM encodings. The default, and when to send wav.

If you send no output_format, Sume's text to speech API returns an mp3 at 44100 Hz and 128 kbps. Send wav with pcm_s16le at 44100 Hz when the audio will feed an avatar mux or when you want sentence slices. The container can be mp3, wav or raw; the sample rate is one of 8000, 16000, 22050, 24000, 44100 or 48000 Hz.
All values are from the TTS request schema in the Sume API reference, read 2026-09-29, and apply to TTS 1.0 and the TTS Router alike.
What can output_format contain?
output_format is an optional object with four fields. Each is an enum in the schema.
| Field | Allowed values | Note in the schema |
|---|---|---|
container | mp3, wav, raw | Audio container |
sample_rate | 8000, 16000, 22050, 24000, 44100, 48000 | Sample rate in Hz |
bit_rate | 32000, 64000, 96000, 128000, 192000 | MP3 bit rate; required for mp3 when overriding defaults |
encoding | pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw | PCM encoding for wav and raw containers |
When should I send wav instead of mp3?
The schema names two cases. The first is feeding the audio into an avatar mux: pass wav, pcm_s16le and 44100 explicitly, because the default is mp3. The second is sentence segmentation. segmentation.mode=sentence needs timestamps.words: true, and each segment gets a sample-exact audio_url only when the container is wav or raw. With mp3 you still get the timings but no slice URLs.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-wav-001" \
-d '{
"transcript": "Your order ships tomorrow. Tracking follows by email.",
"avatar_handle": "acme",
"language": "en",
"output_format": {
"container": "wav",
"sample_rate": 44100,
"encoding": "pcm_s16le"
},
"timestamps": { "words": true },
"segmentation": { "mode": "sentence" }
}'Which combination should I send?
Nothing in the schema ranks the sample rates, so choose by where the file goes. These rows only restate what the schema and its descriptions say.
| Goal | Send | Why the schema says so |
|---|---|---|
| Just get audio back | Omit output_format | Defaults to mp3 44100 Hz, 128 kbps |
| Feed an avatar mux | wav, pcm_s16le, 44100 | The schema says to pass these explicitly for the avatar mux |
| Sentence audio slices | wav or raw, with timestamps.words and segmentation.mode=sentence | Segment audio_url slices come only from wav or raw |
| Pick an mp3 bit rate | mp3 with a bit_rate from 32000 to 192000 | bit_rate is the mp3 setting |
What about telephony sample rates?
The lowest rate on the list is 8000 Hz, and the encodings include pcm_mulaw and pcm_alaw. The schema does not say what any phone system expects, so check your provider's audio spec before you pick one; the docs promise only that these values are accepted.
What limits still apply?
- The
output_formatschema setsadditionalProperties: false, so send only the four fields above. - On TTS 1.0, synthesized audio longer than 1200 seconds fails with
tts_duration_exceeded, whatever the format. - A
bit_rateis for mp3. The schema describesencodingas the PCM setting for wav and raw. - For the endpoints and their prices, see Text to speech API.
- For the word timings themselves, see Text to speech with highlighted words.
Sources
Related posts
More in Models
- Veo 3 shutdown: which Gemini API models retired and what replaces them
Google retired veo-3.0-generate-001, veo-3.0-fast-generate-001 and veo-2.0-generate-001 on 2026-06-30. Its listed replacements are the Veo 3.1 preview models.
- Veo 3.1 4K API: 8 seconds only, and 4K options on Sume
Veo 3.1 makes 4K video only at 8 seconds, and not on the Lite model. On Sume, gemini-omni-flash-1.1 lists 4K for 3 to 10 seconds.
- Veo 3.1 API: how to call it, and what Sume lists instead
The Veo 3.1 API is Google's, called through the Gemini API with preview model codes. Sume does not list Veo; here are the catalog models with the same inputs.
- Veo 3.1 extend video, and how to chain clips on Sume
Veo 3.1 extends only videos Veo made, 7 seconds at a time. Sume has no extend call, so you chain clips from a last frame or pick a model with a longer limit.
Written by Sume