Alexa skill audio from Sume music: 48 kbps MPEG-2 re-encode step

Alexa SSML audio wants MPEG-2 mp3 at 48 kbps and 16000, 22050 or 24000 Hz. Sume documents no music bitrate, so probe the file and re-encode it yourself.

4 min readSume
All posts

An Alexa skill cannot play a Sume music file as it is. Amazon's SSML reference asks for an MPEG-2 mp3 at 48 kbps with a sample rate of 16000, 22050 or 24000 Hz, hosted on HTTPS. Sume's docs say the router audio is usually audio/mpeg but do not state a bitrate, so probe the file with ffprobe, then re-encode it with ffmpeg in your own build step.

What Alexa documents

The audio tag in SSML has several limits at once.

Alexa SSML audio facts (Alexa Skills Kit, read 2026-10-05)
TopicAlexa documents
FormatMPEG-2 mp3, 48 kbps
Sample rates16000, 22050 or 24000 Hz
Total length in outputSpeech240 seconds
Total length in reprompt90 seconds
Files per responseAt most 5
HostingHTTPS with a trusted certificate

What Sume gives you

A Music Router job returns an audio artifact, usually audio/mpeg, on media.sume.com, for a fixed $0.125. Sume's audio detach job takes a video and returns its audio, and Timeline audio joins or splits Sume-hosted audio; neither is documented as a re-encoder to a target bitrate. The conversion belongs to you.

Probe, then re-encode

Run ffprobe to see what you have. Then convert to mono 24 kHz at 48 kbps; the -ar and -b:a flags set the rate and bitrate. Check the result with ffprobe again, because the MPEG version of the output matters: at 24000 Hz or below, the MP3 encoder writes MPEG-2 frames.

ffprobe -v error -show_entries stream=codec_name,sample_rate,channels,bit_rate \
  -of default=nw=1 sume-track.mp3

ffmpeg -y -i sume-track.mp3 -ac 1 -ar 24000 -b:a 48k -codec:a libmp3lame alexa-track.mp3

ffprobe -v error -show_entries stream=codec_name,sample_rate,bit_rate \
  -of default=nw=1 alexa-track.mp3

Trim to the budget

A response may hold 240 seconds of audio in total in outputSpeech, and 90 seconds in a reprompt. Since the router has no duration field, put the length in the prompt and trim with Timeline audio split if needed. For a reprompt, a short sting is the sensible size.

Always test on a real device or the developer console, since these are Amazon's documented limits and your file must also be reachable by Amazon's servers with a trusted certificate.

Before you ship

If you also need loudness or silence trimmed, do it in the same ffmpeg step, so the file is touched only once before you host it.

  • Run ffprobe on the Sume file first and write down codec, rate and bitrate.
  • Re-encode to the Alexa profile with your own ffmpeg step.
  • Host the result on a trusted HTTPS endpoint.
  • Keep the total audio per response under the documented seconds.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume