Audio detach size for a 180 s Short: 34.56 MB wav, 2.88 MB mp3
The byte arithmetic for a 3-minute track from audio detach: stereo 48 kHz wav, mono 16 kHz wav and 128 kbps mp3, with the fields that set each size.

A 180-second track from Sume audio detach is about 34.56 MB as a stereo 48 kHz wav, 17.28 MB as mono 48 kHz, 5.76 MB as the 16 kHz mono wav used for speech-to-text, and 2.88 MB as an mp3. These are arithmetic from the format parameters in the docs, not measurements of a particular file, and they ignore the few dozen bytes of header.
Where the numbers come from
The default output is wav, sample-exact pcm_s16le: 16 bits, or 2 bytes, per sample per channel. Bytes per second is sample_rate * channels * 2. The mp3 option is fixed at 128 kbps, which is 16,000 bytes per second. Multiply by 180 for a three-minute Short.
| Request | Bytes per second | 180 s size |
|---|---|---|
| wav, source channels, 48000 Hz stereo | 192,000 | 34,560,000 B (34.56 MB) |
| wav, 44100 Hz stereo | 176,400 | 31,752,000 B (31.75 MB) |
wav, channels: mono, 48000 Hz | 96,000 | 17,280,000 B (17.28 MB) |
wav, channels: mono, 16000 Hz | 32,000 | 5,760,000 B (5.76 MB) |
| mp3, 128 kbps | 16,000 | 2,880,000 B (2.88 MB) |
Which fields set which row
channels is source (the default) or mono. sample_rate is 16000, 44100 or 48000, and when you omit it the file inherits the source rate; the result reports sample_rate as null in that case, so read the probe if you need the real number. format is wav or mp3. The pair sample_rate: 16000 with channels: mono is the one the docs name as the speech-to-text shape.
Choosing between them
Pick wav when the file will be joined again or will drive lip-sync. The timeline audio page explains that mp3 adds priming padding again at every edge, so sentences cut from an mp3 can pick up small gaps at the joins. Pick mp3 when you only need a listenable copy, since it is roughly one twelfth the size of a 48 kHz stereo wav (34.56 divided by 2.88). The wav is also the format that timeline_create audio.url and POST /v1/timeline-1.0/audio expect.
Limits and errors
A whole track longer than 900 seconds needs a range, because the output cap is 900 s even though the source cap is 1800 s. A clip with no audio track fails with detach_source_has_no_audio, so probe has_audio first with a frames: false inspect. Each detach job is $0.01 per the docs, and the live rate is in GET /v1/catalog.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/short.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000
}Sources
Related posts
More in Media tools
- Caption a talking-head video: style, placement and phrasing
Caption a talking-head clip with the Sume API: pick a style, move the line off the face, set words per card. One request, $0.20 for up to 60 seconds.
- Captions for a video with loud background music: script text or cues
Music can bury speech and trip speech-to-text. Sume documents three ways to supply your own wording: script_text, a words array or cues. Which to pick.
- ChatGPT Sora clip as reference video: Omni takes 3 s each, trim
Sora 2 left the API but a ChatGPT clip can still seed a new render. Gemini Omni Flash 1.1 on Sume takes up to 3 reference videos, each 3 s at most.
- Crop a 16:10 MacBook screen recording to 9:16: width 0.3516
A 16:10 recording needs crop width 0.3515625 for 9:16: 900x1600 at 2560x1600, but 674x1200 at 1920x1200 because 675 is odd. The request and the rounding.
Written by Sume