TTS mp3 bit_rate 32k to 192k: how big is one minute of voiceover?
Sume TTS mp3 bit_rate takes 32000, 64000, 96000, 128000 or 192000, default 128000. At a constant bit rate one minute is 240 KB, 480, 720, 960 or 1,440 KB.

Sume TTS output_format with container: "mp3" accepts bit_rate values of 32000, 64000, 96000, 128000 and 192000, and the default is 128000 at 44100 Hz. At a constant bit rate, bytes per second are the bit rate divided by eight, so one minute is 240 KB at 32k, 480 KB at 64k, 720 KB at 96k, 960 KB at 128k and 1,440 KB at 192k.
The table
These are decimal kilobytes and megabytes, and they ignore small tag overhead. Your file lands within a few percent of the figure.
| bit_rate | Bytes per second | One minute | Ten minutes |
|---|---|---|---|
| 32000 | 4,000 | 240 KB | 2.4 MB |
| 64000 | 8,000 | 480 KB | 4.8 MB |
| 96000 | 12,000 | 720 KB | 7.2 MB |
| 128000 (default) | 16,000 | 960 KB | 9.6 MB |
| 192000 | 24,000 | 1,440 KB | 14.4 MB |
Where the number matters
Size decides whether a file passes an upload limit, loads fast on a phone, or fits a message. A limit that is known on Sume's side is the 10 MB audio ceiling on Fabric lip sync: ten minutes at the default would be 9.6 MB, but a lip-sync clip is far shorter. For a podcast, the default 128000 matches common expectations; for a voice prompt in an app, 64000 is often enough for speech.
Voice is narrow-band content. Lower bit rates sound acceptable for speech where music would suffer, but listen to a take at the rate you plan to ship.
The request
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-mp3-64k-001" \
-d '{
"transcript": "Welcome back. Today we are looking at three small changes.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en",
"output_format": { "container": "mp3", "sample_rate": 44100, "bit_rate": 64000 }
}'Choices that cost nothing in price
The bit rate does not change the bill: TTS is $0.0475 per 1,000 characters regardless. It only changes the file. If the audio will be cut, joined, or used to drive a lip-sync render, ask for wav instead, since mp3 adds padding at the edges when it is encoded. Gap between joined files explains why.
The listed sample rates are 8000, 16000, 22050, 24000, 44100 and 48000. A lower sample rate shrinks the file too, but test it on your voice before you rely on it.
A rule of thumb
Use the default 128000 unless you have a reason to change it. Go down to 64000 for spoken-only app audio where size matters, and up to 192000 only when the audio is mixed with music in a file you will not re-encode. Each step changes only the file, never the TTS price.
Keep the bit rate in your preset so every episode matches.
Related posts
More in Developers
- TTS raw PCM output: container raw needs sample_rate and encoding
Sume TTS can return mp3, wav or headerless raw audio. Raw is the only container where sample_rate and encoding are required, and the fields are strict.
- TTS Router 400 unknown model: which Sonic ids does Sume accept?
POST /v1/tts-router/generate needs a catalog model id; an unknown one returns 400 with catalog_url. Seed ids: sonic-3.6, 3.5, 3, latest, preview. Curl inside.
- Korean TTS segment text has no spaces, unless a digit is in it
Sume's TTS segment text joins tokens with spaces only if one has a Latin letter or digit; else with nothing. Use segments for timing, your script for text.
- TTS sentence slices: mp3 gives timings only, wav gives audio_urls
Sume TTS segmentation returns sentence timings for any container, but slice audio_urls only with wav or raw. Request shape, the 70 ms rule and when to pick wav.
Written by Sume