TTS mp3 bit_rate vs wav: fit a voiceover under the 10 MB Fabric limit
Sume TTS mp3 bit rates run 32k to 192k. At 128k a 300-second voiceover is about 4.8 MB, under the 10 MB Fabric audio limit. Mono 16 kHz wav is 9.6 MB.

Choose mp3 for audio that goes into Fabric: at the default 128 kbps a 300-second voiceover is about 4.8 MB, well under the 10 MB audio_url limit, while 16-bit mono wav at 16 kHz is about 9.6 MB and wav at 48 kHz mono is about 28.8 MB. The arithmetic is bytes per second times seconds, and it tells you the format before you hit a size error.
The wav figures assume a single channel; stereo doubles them.
The numbers in the contract
The TTS contract sets output_format defaults to mp3 at 44100 Hz and 128000 bits per second. The mp3 bit rate options are 32000, 64000, 96000, 128000 and 192000. Containers are mp3, wav and raw, sample rates run 8000, 16000, 22050, 24000, 44100 and 48000, and PCM encodings include pcm_s16le, pcm_f32le, pcm_mulaw and pcm_alaw. The Fabric route takes a Sume-hosted audio_url of at most 10 MB and a duration_seconds from 1 to 300, as the OpenAPI describes.
The ElevenLabs text-to-speech doc lists MP3, PCM, mu-law, A-law and Opus formats; Opus is not in the Sume TTS list.
The arithmetic
A bit rate in bits per second divided by 8 gives bytes per second. Multiply by the duration. For 16-bit PCM the bytes per second are sample rate times 2 per channel.
| Format | Bytes per second | 300 seconds |
|---|---|---|
| mp3 32 kbps | 4,000 | 1.2 MB |
| mp3 64 kbps | 8,000 | 2.4 MB |
| mp3 128 kbps | 16,000 | 4.8 MB |
| mp3 192 kbps | 24,000 | 7.2 MB |
| wav 16-bit, 16 kHz | 32,000 | 9.6 MB |
| wav 16-bit, 48 kHz | 96,000 | 28.8 MB |
What to pick
For Fabric, the file must be small enough and the sound clean enough to drive the mouth, so mp3 at 128 kbps is a safe default. For editing, keep wav, since each mp3 pass adds padding at the edges and the timeline audio docs recommend wav for audio that will be joined again or that drives lip-sync. A reasonable pipeline makes wav, edits, and exports an mp3 copy for the Fabric upload.
- Voice only: 64 to 128 kbps is plenty for speech.
- Speech to text: use 16 kHz mono, as in the detach post.
- Joining later: stay in wav until the end.
- Phone-style delivery: the mulaw and alaw encodings with 8000 Hz exist for that.
Check before you send
Read the file size from the artifact before submitting, and compute duration from the job result. If you are close to 10 MB, drop the bit rate one step rather than trimming content. Remember that file size is not the only gate: Fabric also needs duration_seconds between 1 and 300.
Sources
Related posts
More in Developers
- TTS word timings to burned-in captions: send them as words on Sume
Sume's TTS can return word start and end times; the caption job accepts words with text, start and end and skips transcription. How to wire them together.
- Two workers, one order: Idempotency-Key from order id and version
Two queue workers pick up the same order and both submit to Sume. Build the key from order id plus version so duplicates collapse and edits still create a run.
- Undo for AI image edits: keep a version chain of every saved result
Generative edits have no undo button. Download each result, hash it, record its parent and prompt, and walk the chain back. Python, no database.
- unsupported_media_type: video_url served as text/html or an image
Sume video trim and filter HEAD the source and refuse a declared non-video content type. What is checked, why octet-stream passes, and a runnable check.
Written by Sume