A 5-minute voice-over: 4.8 MB as 128 kbps MP3, 26.5 MB as WAV
Sume TTS defaults to MP3 at 44,100 Hz and 128 kbps. Five minutes is 4.8 MB; 16-bit mono WAV is 26.46 MB at 44.1 kHz and 28.8 MB at 48 kHz.

A five-minute voice-over is 4.8 MB as a default Sume TTS MP3 (128 kbps) and 26.46 MB as 16-bit mono WAV at 44.1 kHz, or 28.8 MB at 48 kHz. The WAV sizes assume one channel and 16-bit samples, which is my assumption; stereo doubles them. Choose MP3 when the file goes straight to a listener, and WAV or raw when it goes on to another processing step such as cutting, joining or sentence-level timing.
The arithmetic
MP3 at 128 kbps is 16,000 bytes a second, so 300 seconds is 4,800,000 bytes. Uncompressed audio is sample rate times bytes per sample times channels: 44,100 x 2 x 1 = 88,200 bytes a second, which is 26,460,000 bytes for 300 seconds. At 48,000 Hz the rate is 96,000 bytes a second and 28,800,000 bytes.
| Format | Per second | Five minutes | Notes |
|---|---|---|---|
| MP3, 44,100 Hz, 128 kbps (default) | 16,000 bytes | 4.8 MB | Default output_format |
| WAV, 16-bit mono, 44,100 Hz | 88,200 bytes | 26.46 MB | Assumes mono and 16 bit |
| WAV, 16-bit mono, 48,000 Hz | 96,000 bytes | 28.8 MB | Sample rates 8,000 to 48,000 are accepted |
| WAV, 16-bit mono, 16,000 Hz | 32,000 bytes | 9.6 MB | Common speech-to-text rate |
Which step needs which
The TTS docs tie sentence timing to the output: emit_audio for sentence segmentation needs wav or raw output. Cutting and joining audio with Timeline audio is cleaner on WAV, because MP3 frames add a little padding at every edge. For a mux onto video, either works. For an upload to a platform or a download link, MP3 is a fifth of the size.
Speech to text accepts an audio_url with its own size limit, so check the OpenAPI reference before you send long WAV files to transcription; the MP3 is a fifth of the size.
A default that holds up
Keep the master as WAV for cutting, create an MP3 for delivery, and measure the real file before you assume a size. Both cost the same TTS price; the format does not change the $0.0475 per 1,000 characters.
Rates other than 44.1 kHz
Output sample rates run from 8,000 to 48,000 Hz. A lower rate shrinks WAV in direct proportion: 8,000 Hz mono 16-bit is 16,000 bytes a second, 4.8 MB for five minutes, the same as the default MP3, but with the quality of a telephone line. Speech models usually want 16,000 Hz, so a 16 kHz mono WAV is a reasonable working copy at 9.6 MB. The size of the file never changes the speech price.
Sources
Related posts
More in Developers
- TTS sentence slices: segmentation, wav output and boundary_lead_ms 70
Sume TTS can return gapless sentence segments. Needs timestamps.words and wav or raw for per-sentence audio_url. A 900-character script costs $0.04275.
- Sume TTS sentence segments: boundary_lead_ms 0 vs 500, 12 lines
Ask Sume TTS for words and sentence segmentation and you get gapless wav slices. boundary_lead_ms (default 70) moves each cut 0 to 500 ms after the last word.
- TTS speed 1.2 turns a 30-second read into about 25 seconds
Sume TTS accepts generation_config.speed from 0.6 to 1.5. Speed changes length, not the bill: 450 characters cost $0.021375 at any speed. Test the real length.
- TTS word timings straight into captions: a 45-second ad for 33 cents
Ask Sume TTS for timestamps.words and send them as words on the caption job: no speech-to-text. 700 characters, render and captions come to $0.33.
Written by Sume