TTS mp3 bit_rate vs wav: fit a voiceover under the 10 MB Fabric limit

Sume TTS mp3 bit rates run 32k to 192k. At 128k a 300-second voiceover is about 4.8 MB, under the 10 MB Fabric audio limit. Mono 16 kHz wav is 9.6 MB.

4 min readSume
All posts

Choose mp3 for audio that goes into Fabric: at the default 128 kbps a 300-second voiceover is about 4.8 MB, well under the 10 MB audio_url limit, while 16-bit mono wav at 16 kHz is about 9.6 MB and wav at 48 kHz mono is about 28.8 MB. The arithmetic is bytes per second times seconds, and it tells you the format before you hit a size error.

The wav figures assume a single channel; stereo doubles them.

The numbers in the contract

The TTS contract sets output_format defaults to mp3 at 44100 Hz and 128000 bits per second. The mp3 bit rate options are 32000, 64000, 96000, 128000 and 192000. Containers are mp3, wav and raw, sample rates run 8000, 16000, 22050, 24000, 44100 and 48000, and PCM encodings include pcm_s16le, pcm_f32le, pcm_mulaw and pcm_alaw. The Fabric route takes a Sume-hosted audio_url of at most 10 MB and a duration_seconds from 1 to 300, as the OpenAPI describes.

The ElevenLabs text-to-speech doc lists MP3, PCM, mu-law, A-law and Opus formats; Opus is not in the Sume TTS list.

The arithmetic

A bit rate in bits per second divided by 8 gives bytes per second. Multiply by the duration. For 16-bit PCM the bytes per second are sample rate times 2 per channel.

Size of a 300-second voiceover by format, single channel (read 2026-10-03)
FormatBytes per second300 seconds
mp3 32 kbps4,0001.2 MB
mp3 64 kbps8,0002.4 MB
mp3 128 kbps16,0004.8 MB
mp3 192 kbps24,0007.2 MB
wav 16-bit, 16 kHz32,0009.6 MB
wav 16-bit, 48 kHz96,00028.8 MB

What to pick

For Fabric, the file must be small enough and the sound clean enough to drive the mouth, so mp3 at 128 kbps is a safe default. For editing, keep wav, since each mp3 pass adds padding at the edges and the timeline audio docs recommend wav for audio that will be joined again or that drives lip-sync. A reasonable pipeline makes wav, edits, and exports an mp3 copy for the Fabric upload.

  • Voice only: 64 to 128 kbps is plenty for speech.
  • Speech to text: use 16 kHz mono, as in the detach post.
  • Joining later: stay in wav until the end.
  • Phone-style delivery: the mulaw and alaw encodings with 8000 Hz exist for that.

Check before you send

Read the file size from the artifact before submitting, and compute duration from the job result. If you are close to 10 MB, drop the bit rate one step rather than trimming content. Remember that file size is not the only gate: Fabric also needs duration_seconds between 1 and 300.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume