wav or mp3 for joined audio: priming padding at every edge
Sume timeline audio outputs wav by default and mp3 as an option. Keep wav if you join again or lip-sync: mp3 adds priming padding at every edge.

Keep wav. Sume timeline audio's output.format is wav by default (pcm_s16le, sample-exact) or mp3, which is smaller but adds priming padding again at every edge. The docs say to keep wav if you will join the file again or if the file drives lip-sync. Use mp3 only for a final file where size matters more than sample accuracy.
The two formats side by side
| Format | Encoding | Edges | Use for |
|---|---|---|---|
wav (default) | pcm_s16le | Sample-exact | Re-joins, lip-sync, render inputs |
mp3 | Compressed | Priming padding added again at every edge | Final delivery files |
Why edges matter
A join is only gapless if the seams are exact. Timeline audio joins in the sample domain, with no re-TTS and no silence at the seams. With mp3 output, every edge gets priming padding, and joining those files again stacks the padding. Cut four times and the padding appears four times.
Lip-sync is the strict case. Avatar 1.0 image-to-video audio and Timeline 1.0 audio.url can both take the joined file, and a drift of a few milliseconds is what a face shows first. Detach also defaults to wav for the same reason; its mp3 option is 128 kbps.
Limits that apply to both
The produced audio is up to 1800 seconds. A concat takes 1 to 20 parts, and a split takes 1 to 20 ranges that may overlap. The public rate is $0.01 flat per job and uses no provider inference, only worker ffmpeg. Confirm the live rate in GET /v1/catalog. Format choice does not change the price in the docs.
- Chain wav to wav, convert once at the end.
- If you must send mp3 into a later join, expect the padding.
- Import files first with
POST /v1/media-imports.
Request with mp3 output
Send output.format only when you really want mp3.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timeline-audio-mp3-001" \
-d '{
"operation": "concat",
"output": { "format": "mp3" },
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/a.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/b.wav" }
]
}'Sources
Related posts
More in Media tools
- Which AI video models make a 20-second clip in one job on Sume?
Seedance 2.5 (4-30 s) and Wan 3.0 (2-30 s) cover 20 seconds in one job on Sume. Omni Flash 1.1 stops at 10 s, MiniMax H3 at 15 s. Full duration table.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
Written by Sume