Amazon audio ads, 3 MB cap: 30 s mono WAV fits, stereo does not
Amazon audio ads allow 10 to 30 seconds and 3 MB. In 16-bit PCM, 30 s mono at 44.1 kHz is 2.65 MB and stereo is 5.29 MB. Detach it with Sume audio-detach.

A 30-second WAV at 44.1 kHz in mono is about 2.65 MB and fits Amazon's 3 MB audio ad cap, while the same clip in stereo is about 5.29 MB and does not. If your audio comes out of a video, Sume's audio-detach can return exactly that mono WAV for $0.01 per job.
Amazon's audio ad page (read 2026-10-04) lists a length of 10 to 30 seconds, a maximum of 3 MB, formats WAV, MP3 or OGG, and audio quality of at least 192 kbps with RMS normalized to -14 dBFS and peak normalized to -0.2 dBFS. It adds a required 1024x1024 companion banner. The page does not state sample rate or channel count, which is why the arithmetic below matters.
How big is a WAV at each setting?
Audio-detach writes WAV as 16-bit signed PCM, so size is sample rate times 2 bytes times channels times seconds. The figures below use decimal megabytes and a 30-second clip.
| Sample rate | Channels | Size | Under 3 MB? |
|---|---|---|---|
| 16000 Hz | Mono | 0.96 MB | Yes |
| 44100 Hz | Mono | 2.65 MB | Yes |
| 48000 Hz | Mono | 2.88 MB | Yes, tight |
| 44100 Hz | Stereo | 5.29 MB | No |
| 48000 Hz | Stereo | 5.76 MB | No |
Why not use MP3 instead?
Audio-detach's MP3 option is fixed at 128 kbps, and Amazon asks for at least 192 kbps. A 128 kbps file would miss the stated quality floor even though it is far smaller. WAV is the safer choice, and mono at 44.1 kHz is the largest setting that clears the cap with room to spare.
What does the Sume call look like?
Audio-detach takes a media.sume.com video, an optional range, format, channels (source or mono) and a sample_rate of 16000, 44100 or 48000. The source can be up to 1800 seconds and the output up to 900. Each request needs an Idempotency-Key header. Trim to your 30 seconds first with video-trim, or pass a range, then detach.
import os, uuid, requests
resp = requests.post(
"https://api.sume.com/v1/audio-detach",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"video_url": os.environ["SUME_SOURCE_URL"],
"mode": "sync",
"format": "wav",
"channels": "mono",
"sample_rate": 44100,
},
timeout=60,
)
print(resp.status_code, resp.json())What does Sume not do for me here?
The audio-detach docs describe no loudness operation, so the -14 dBFS RMS and -0.2 dBFS peak targets are yours to hit and measure elsewhere. Check the duration in the response too: Amazon wants 10 to 30 seconds, and a 30.4-second file would be out of range.
For a related loudness target on a video ad, see the stored post on Spotify video ad specs, and for the audio ad basics see audio ad specs.
Sources
Related posts
More in Media tools
- Burn feature text onto silent product clips with caption cues
A silent product clip fails speech captions. Pass your own cues (text, start, end) to Sume's video-captions endpoint and burn feature text on for TikTok Shop.
- Caption a silent AI video: fixing caption_no_speech
A silent clip fails POST /v1/video-captions with caption_no_speech. Send cues with text, start and end to burn authored captions without speech-to-text.
- Caption a silent Seedance 2.5 clip with authored cues
A clip with no speech fails Sume caption jobs as caption_no_speech. Pass cues with text, start and end seconds instead; $0.20 per job up to 60 seconds.
- Change the caption highlight colour without changing the style
Cheap transcription made captions routine; brand colour is what is left. One design field changes the spoken-word colour and keeps everything else in the style.
Written by Sume