Spotify audio ad specs: length, file, and loudness rules

Spotify audio ads run up to 30 s as MP3, WAV or OGG: 44.1 kHz stereo, 192–320 kbps, -16 LUFS, up to 50 MB, plus a 600×600-minimum companion.

5 min readSume
All posts

Spotify audio ad specs: a music ad runs up to 30 seconds, and the file is MP3, WAV or OGG, up to 50 MB, at 44.1 kHz, 192 to 320 kbps, stereo, and -16 LUFS (±1.5 LUFS). Spotify recommends no more than 65 words for a 30-second ad, and the ad carries a 1:1 companion image of at least 600×600 pixels.

Spotify's numbers are quoted from its Audio ad specs & requirements page for Ads Manager and Direct IO, read on 2026-09-29; Spotify can change them. Sume facts come from the Timeline 1.0 and Audio detach docs and the TTS schema in the Sume API reference. Anything called current behavior is read from Sume's code.

What are Spotify's audio ad file requirements?

The technical rules are the same for music and podcast ads bought through Ads Manager or Direct IO:

From Spotify's Audio ad specs & requirements, read 2026-09-29.
SpecSpotify's rule
LengthMusic ads: 30 s maximum; podcast ads: 30 s maximum
File typeMP3, WAV or OGG
File size50 MB maximum
Sample rate44.1 kHz
Bit rate192 kbps minimum, 320 kbps maximum
Loudness-16 LUFS (±1.5 LUFS)
ChannelsStereo
ScriptNo more than 65 words for 30 s; 100 words for 60 s

Can a Spotify audio ad be 60 seconds?

Only in some cases. Spotify says long-form music ads are available in select markets for Direct campaigns, can run up to 60 seconds, and are skippable after 30 seconds. Apart from the length, they follow the same specs as a standard audio ad. For a 60-second read, Spotify recommends a script of 100 words.

What else goes with the audio file?

A music ad is more than the sound. Spotify lists these creative pieces:

  • Companion display unit: 1:1, a minimum of 600×600 pixels, 200 KB maximum, JPEG or PNG. It shows when listeners engage with the app during the break.
  • Advertiser name of no more than 25 characters, and one call-to-action from Spotify's list, such as “Learn more” or “Shop now”.
  • Tagline of 40 characters maximum, and one clickthrough URL on HTTPS.
  • Optional Canvas: a 9:16 loop at 720×1280, MP4 or MOV, 3–8 seconds, 200 KB maximum, which replaces the companion on mobile.

How do I make a Spotify audio ad with Sume?

For a voice-only spot, text to speech can write Spotify's sample rate and bitrate directly: POST /v1/tts-1.0/generate takes an output_format with container mp3, sample_rate 44100, and bit_rate 192000 (the default is 128 kbps, under Spotify's floor). It costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default. The schema does not state a channel count, so check that the file is stereo.

For voice over a music bed, the mix is a Timeline 1.0 render, then audio detach; AI radio commercial generator walks through both jobs. Detach can set sample_rate to 44100, but its two formats miss Spotify's bitrate window: the MP3 is 128 kbps, and a 44.1 kHz, 16-bit stereo WAV runs at 1,411.2 kbps, above the 320 kbps maximum Spotify lists. Encode the final MP3 or OGG at 192–320 kbps in another tool.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: spotify-spot-001" \
  -d '{
    "transcript": "Your 65-word script goes here.",
    "voice": { "mode": "id", "id": "your-voice-id" },
    "output_format": { "container": "mp3", "sample_rate": 44100, "bit_rate": 192000 },
    "mode": "async"
  }'

What does Sume leave you to check?

Loudness and stereo are not settings in Sume, and the length is yours to set:

  • Loudness: Sume's Timeline and audio detach docs list no loudness setting. Timeline's level controls are audio.gain_db (−60 to 12) and the soundtrack's gain_db and duck_db, so measure and set loudness in an audio tool.
  • Stereo: audio detach's channels is source or mono, so it cannot turn a mono voice into stereo.
  • Length: a Timeline render's length is its audio.duration_seconds, so set it to 30 or less. WAV vs MP3 explains why to keep the WAV until that last encode.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume