Alexa skill audio from Sume music: 48 kbps MPEG-2 re-encode step
Alexa SSML audio wants MPEG-2 mp3 at 48 kbps and 16000, 22050 or 24000 Hz. Sume documents no music bitrate, so probe the file and re-encode it yourself.

An Alexa skill cannot play a Sume music file as it is. Amazon's SSML reference asks for an MPEG-2 mp3 at 48 kbps with a sample rate of 16000, 22050 or 24000 Hz, hosted on HTTPS. Sume's docs say the router audio is usually audio/mpeg but do not state a bitrate, so probe the file with ffprobe, then re-encode it with ffmpeg in your own build step.
What Alexa documents
The audio tag in SSML has several limits at once.
| Topic | Alexa documents |
|---|---|
| Format | MPEG-2 mp3, 48 kbps |
| Sample rates | 16000, 22050 or 24000 Hz |
| Total length in outputSpeech | 240 seconds |
| Total length in reprompt | 90 seconds |
| Files per response | At most 5 |
| Hosting | HTTPS with a trusted certificate |
What Sume gives you
A Music Router job returns an audio artifact, usually audio/mpeg, on media.sume.com, for a fixed $0.125. Sume's audio detach job takes a video and returns its audio, and Timeline audio joins or splits Sume-hosted audio; neither is documented as a re-encoder to a target bitrate. The conversion belongs to you.
Probe, then re-encode
Run ffprobe to see what you have. Then convert to mono 24 kHz at 48 kbps; the -ar and -b:a flags set the rate and bitrate. Check the result with ffprobe again, because the MPEG version of the output matters: at 24000 Hz or below, the MP3 encoder writes MPEG-2 frames.
ffprobe -v error -show_entries stream=codec_name,sample_rate,channels,bit_rate \
-of default=nw=1 sume-track.mp3
ffmpeg -y -i sume-track.mp3 -ac 1 -ar 24000 -b:a 48k -codec:a libmp3lame alexa-track.mp3
ffprobe -v error -show_entries stream=codec_name,sample_rate,bit_rate \
-of default=nw=1 alexa-track.mp3Trim to the budget
A response may hold 240 seconds of audio in total in outputSpeech, and 90 seconds in a reprompt. Since the router has no duration field, put the length in the prompt and trim with Timeline audio split if needed. For a reprompt, a short sting is the sensible size.
Always test on a real device or the developer console, since these are Amazon's documented limits and your file must also be reachable by Amazon's servers with a trusted certificate.
Before you ship
If you also need loudness or silence trimmed, do it in the same ffmpeg step, so the file is touched only once before you host it.
- Run ffprobe on the Sume file first and write down codec, rate and bitrate.
- Re-encode to the Alexa profile with your own ffmpeg step.
- Host the result on a trusted HTTPS endpoint.
- Keep the total audio per response under the documented seconds.
Sources
Related posts
More in Media tools
- Amazon Sponsored Brands video audio: PCM, AAC or MP3, and Sume
Amazon Sponsored Brands video takes PCM, AAC or MP3 audio, 16:9 only, 6-45 s. Sume's audio-detach gives wav pcm_s16le or 128 kbps mp3; a Sume MP4 needs a probe.
- Animate an infographic with Wan 3.0: first frame or reference image?
Use the infographic as a first frame to keep the layout, or as a reference to get new motion. How Sume decides, and which to pick for charts, icons and text.
- Assert an edited image has the source's pixel size: Python check
A short Python check that downloads the source and the edited file and fails if the pixel size differs. Needed on Sume whenever you send an aspect_ratio.
- Audio detach range without end: take the rest of a track
In Sume audio detach, range.end is optional: a start-only range runs to the end of the track. How it meets the 900 s output cap and 1800 s source cap.
Written by Sume