Audio detach file size: 16 kHz mono WAV vs 48 kHz stereo

A 900-second Sume audio detach is about 29 MB as 16 kHz mono WAV, 173 MB as 48 kHz stereo WAV and 14 MB as 128 kbps MP3. The arithmetic and choices.

5 min readSume
All posts

How big is a detached audio track? Using plain arithmetic on the options in the Sume audio detach docs, a full 900-second output is about 28.8 MB as 16 kHz mono WAV, about 172.8 MB as 48 kHz stereo WAV, and about 14.4 MB as 128 kbps MP3.

These are calculated sizes, not measurements. They ignore file headers, so a real file is a little larger.

The options

Audio detach returns a new audio artifact from one Sume-hosted video. format is wav (default, pcm_s16le) or mp3 at 128 kbps. channels is source or mono, and sample_rate is 16000, 44100 or 48000, or omitted to inherit the source. The docs call 16000 plus mono the speech-to-text shape.

Caps: the source can be up to 1,800 seconds, and the output up to 900 seconds. A whole track longer than 900 seconds needs a range. The price is $0.01 per job whatever the size.

Calculated size of a 900-second detach by format (read 2026-10-03)
FormatBytes per secondPer minute900 seconds
WAV 16 kHz mono, 16-bit32,0001.92 MB28.8 MB
WAV 44.1 kHz mono, 16-bit88,2005.29 MB79.4 MB
WAV 48 kHz stereo, 16-bit192,00011.52 MB172.8 MB
MP3 128 kbps16,0000.96 MB14.4 MB

How the numbers are built

Uncompressed PCM size is sample rate times bytes per sample times channels. At 16-bit, that is 2 bytes per sample. So 16,000 times 2 times 1 is 32,000 bytes per second, and 900 seconds of that is 28.8 million bytes. MP3 at 128 kbps is 128,000 bits per second, which is 16,000 bytes per second. I use 1 MB as one million bytes.

A stereo source left on channels: source doubles the mono figure. If you only need speech, mono halves the file for free.

Choosing a shape

Match the shape to the next step.

  • Speech to text: 16 kHz mono WAV. It is the smallest uncompressed option and the docs name it as the STT shape.
  • Joining or lip-sync: keep WAV. The timeline audio docs say MP3 re-adds priming padding at every edge, so keep WAV when the file will be joined again.
  • Listening copy or a podcast feed: MP3 at 128 kbps, the smallest of the four.
  • Reference clips for a voice: use a range so you carry only the seconds you need.

Why size matters for the workflow

Large files cost time at each hop: the upload to your own storage, the download by a later job, and any queue limit your own system has. A 173 MB file that only feeds a transcriber is wasted bandwidth. Choosing 16 kHz mono for that step makes the file about six times smaller than the 48 kHz stereo one, with no change in what the transcriber needs.

Detach once and reuse the artifact. The docs say that for many ranges you detach once, then split with timeline audio, so a single $0.01 detach feeds several clips. Jobs return by polling or webhook, as covered in Sume jobs and results.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume