Spotify and Amazon audio ads share one loudness target: -14 RMS
Spotify's video ad page and Amazon's audio ad page both ask for RMS at -14 dBFS and peak at -0.2 dBFS. Where Sume's gain_db helps and where you must measure.

One mix can serve both: Spotify's in-stream video ad page and Amazon's audio ad page each ask for audio with RMS normalized to -14 dBFS and peak normalized to -0.2 dBFS. Sume's timeline has a gain_db field on the audio spine to move the level, but nothing in the docs measures loudness, so check the result with a meter you trust before you upload.
Spotify's spec page (read 2026-10-04) states both figures and says it accepts creative up to 30 seconds. Amazon's audio page (read 2026-10-04) states the same two numbers, plus at least 192 kbps. Neither page says how they measure RMS, so use the same method on every file.
Where do the two specs differ?
The loudness lines match; almost everything else does not.
| Item | Spotify in-stream video | Amazon audio |
|---|---|---|
| RMS | -14 dBFS | -14 dBFS |
| Peak | -0.2 dBFS | -0.2 dBFS |
| Length | Up to 30 seconds | 10 to 30 seconds |
| File size | 500 MB maximum | 3 MB maximum |
| Formats | MOV, MP4 | WAV, MP3, OGG |
| Bitrate | 198 kbps ideal, 320 kbps maximum | At least 192 kbps |
How do I move the level in Sume?
Timeline's audio spine accepts an url or parts, a mode, a source_in offset and a gain_db value. Use a negative gain to pull a hot mix down, and re-render. A timeline render costs $0.10 per output minute, rounded up, and needs an Idempotency-Key. The loudness numbers themselves are yours to measure outside Sume, because the docs list no normalization step.
For the audio-only Amazon deliverable, use audio-detach to pull a WAV from the finished video, as the post on Amazon's 3 MB cap shows. Detach is $0.01 per job and does not change the level.
What is a safe workflow?
Build the mix once, render the video, detach the WAV, and measure both files. If the peak sits above -0.2 dBFS, lower the gain by the overshoot plus a small margin and render again. Lowering gain also lowers RMS, so check both numbers after each change instead of one.
Peak and RMS interact: a dynamic mix can hit the peak ceiling long before its RMS reaches -14 dBFS. That is a mastering problem, not a Sume setting, and a limiter in your audio tool is the usual fix.
Is it worth building one mix for both?
Yes if the same voice and music run on both services, since you pass the check once. If the video has loud effects that Amazon's audio-only cut does not need, give each its own mix. The stored posts on Spotify video ad specs and audio ad specs list the other numbers.
Sources
Related posts
More in Media tools
- Stability AI's Series B and label backers: what it means for audio
Stability released Stable Audio 3.0 on 5/20/26 and raised a Series B on 8/25/26 with EA, Sony, UMG and WMG. A dated timeline and what to verify for video work.
- Stabilize shaky video by API: Sume has no deshake, so do this
Sume's video filter refuses deshake and vidstab as unknown filters. Prove it with the free check, stabilize upstream, then use Sume for trim, crop and captions.
- Temu detail video up to 180 s: stitch 30 s clips with Timeline
Temu detail video allows up to 180 s and 300 MB at 720p or higher. Sume clips top out at 30 s, so six clips stitched with Timeline cover the full length.
- TikTok trending API: browse without a query is flag-gated
Sume trending search needs a query on production. Browse with no query, limit up to 100 and relevance floors are dev-only until a flag is set. The differences.
Written by Sume