Audio detach file size: 16 kHz mono WAV vs 48 kHz stereo
A 900-second Sume audio detach is about 29 MB as 16 kHz mono WAV, 173 MB as 48 kHz stereo WAV and 14 MB as 128 kbps MP3. The arithmetic and choices.

How big is a detached audio track? Using plain arithmetic on the options in the Sume audio detach docs, a full 900-second output is about 28.8 MB as 16 kHz mono WAV, about 172.8 MB as 48 kHz stereo WAV, and about 14.4 MB as 128 kbps MP3.
These are calculated sizes, not measurements. They ignore file headers, so a real file is a little larger.
The options
Audio detach returns a new audio artifact from one Sume-hosted video. format is wav (default, pcm_s16le) or mp3 at 128 kbps. channels is source or mono, and sample_rate is 16000, 44100 or 48000, or omitted to inherit the source. The docs call 16000 plus mono the speech-to-text shape.
Caps: the source can be up to 1,800 seconds, and the output up to 900 seconds. A whole track longer than 900 seconds needs a range. The price is $0.01 per job whatever the size.
| Format | Bytes per second | Per minute | 900 seconds |
|---|---|---|---|
| WAV 16 kHz mono, 16-bit | 32,000 | 1.92 MB | 28.8 MB |
| WAV 44.1 kHz mono, 16-bit | 88,200 | 5.29 MB | 79.4 MB |
| WAV 48 kHz stereo, 16-bit | 192,000 | 11.52 MB | 172.8 MB |
| MP3 128 kbps | 16,000 | 0.96 MB | 14.4 MB |
How the numbers are built
Uncompressed PCM size is sample rate times bytes per sample times channels. At 16-bit, that is 2 bytes per sample. So 16,000 times 2 times 1 is 32,000 bytes per second, and 900 seconds of that is 28.8 million bytes. MP3 at 128 kbps is 128,000 bits per second, which is 16,000 bytes per second. I use 1 MB as one million bytes.
A stereo source left on channels: source doubles the mono figure. If you only need speech, mono halves the file for free.
Choosing a shape
Match the shape to the next step.
- Speech to text: 16 kHz mono WAV. It is the smallest uncompressed option and the docs name it as the STT shape.
- Joining or lip-sync: keep WAV. The timeline audio docs say MP3 re-adds priming padding at every edge, so keep WAV when the file will be joined again.
- Listening copy or a podcast feed: MP3 at 128 kbps, the smallest of the four.
- Reference clips for a voice: use a
rangeso you carry only the seconds you need.
Why size matters for the workflow
Large files cost time at each hop: the upload to your own storage, the download by a later job, and any queue limit your own system has. A 173 MB file that only feeds a transcriber is wasted bandwidth. Choosing 16 kHz mono for that step makes the file about six times smaller than the 48 kHz stereo one, with no change in what the transcriber needs.
Detach once and reuse the artifact. The docs say that for many ranges you detach once, then split with timeline audio, so a single $0.01 detach feeds several clips. Jobs return by polling or webhook, as covered in Sume jobs and results.
Sources
Related posts
More in Developers
- "Avatar does not have a usable TTS voice": the 400 and its fixes
Sume TTS with avatar_id or avatar_handle returns 400 when the avatar has no TTS voice, or when voice.id disagrees with it. What each message means and the fix.
- Try Avatar 1.0 in the Sume playground before you write any code
Use the Sume Avatar playground to validate an avatar or avatar video payload, then move the same body into curl, the CLI or an agent without a rewrite.
- axios rejects Sume 409 and 429: read response.data.error.code
axios rejects any status outside 200-299 by default, so a Sume 409 or 429 throws. Read error.response.data.error.code, and use AbortSignal for timeouts.
- Batch of 30-second Seedance clips: Pro, Startup and Scale queues
How many 30-second Seedance 2.5 jobs can a Sume workspace hold? Plan concurrency and queue limits, what happens at job 25 on Pro, and a safe submit loop.
Written by Sume