Turn a Lyria 3.5 MP3 into a WAV with one Timeline audio job
Sume's music jobs usually return MP3. A single-part Timeline audio concat can output sample-exact WAV for $0.01, which matters before you cut or re-join it.

A Sume Music Router job usually returns an MP3 on media.sume.com. To get a sample-exact WAV, send the track as the only part of a Timeline audio concat. The output defaults to pcm_s16le WAV, for $0.01 flat per job. Do it when the next step is cutting the track, joining it to other audio, or feeding it to a sync-sensitive job.
Why bother
The Timeline audio docs say MP3 output adds encoder priming padding at each edge, while WAV is sample-exact. For a music bed under a voiceover that does not matter. It matters when you cut the track at exact points or join it to other audio, because the padding becomes part of the file and the offsets you computed no longer line up.
Google's music generation page lists Lyria 3.5 as 44.1 kHz stereo, MP3 by default with WAV optional through response_format, and songs of about a couple of minutes controlled by the prompt. The Sume router decides the delivery format, so convert after the job.
The request
Use the artifact URL from the music job result as the single part. The parts array takes 1 to 20 items, so one is valid.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: music-to-wav-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/track.mp3" }
]
}'What you get back
The result is kind: timeline_audio with one audio_url, a duration_seconds, and segments[] giving the offsets. The default output.format is wav. The output may not exceed 1,800 seconds. Request mp3 only if you want a smaller file and accept the padding again.
Replace the demo URL with your real artifact URL from result.artifacts[], where type is audio.
When to skip the step
If the track goes straight under a voiceover in Timeline, use the MP3 as the soundtrack URL. Timeline mixes it for you. Convert to WAV only when the next step is cutting, joining, or a job where timing is checked. Cutting is a separate Timeline audio split call with ranges[], up to 20 per job.
Idempotency and cost
Send an Idempotency-Key so a retry does not run the job twice. The charge is $0.01 flat per job, regardless of track length within the 1,800-second output limit. A music take plus a conversion is therefore $0.135 in total.
If you need to join the track to other audio at the same time, put all the parts in the same concat call and convert once.
Sources
Related posts
More in Media tools
- Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file
A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.
- Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.
- Microsoft's Content Provenance Detection: what you can check on a file
Foundry has a detection website and API for provenance. What the page says it checks, its limits, and how to use it on an AI clip or image from any generator.
- Microsoft: C2PA may not survive a crop, transcode or compression
Microsoft's provenance page lists the edits that can drop credentials and watermarks. Which of them a Sume trim, filter or caption job could be, and a check.
Written by Sume