AI music longer or shorter than the video: prompt, loop, trim or join
A generated track rarely matches the video length. Four ways on Sume: put the length in the prompt, loop the soundtrack, set an in-point, or join takes.

When a music track and a video differ in length, pick the fix by how far apart they are. Ask for the length in the prompt, since Sume rejects a duration field. If the track is a little short, set soundtrack.loop. If it is long, the render's declared length trims it and audio.source_in sets an in-point. If you need much more, join takes with timeline audio concat, up to 20 parts.
The four tools
Each tool belongs to a different surface, so choose by where the fix should live.
| Situation | Tool | Key fields or limits |
|---|---|---|
| Plan ahead | Music Router prompt | Say "a 2-minute track" or use [0:00-0:30] section markers; duration is rejected |
| Track slightly short | Timeline soundtrack | loop, fade_out_seconds up to 10 |
| Track long, want a later start | Timeline audio spine | audio.source_in; output length is still audio.duration_seconds |
| Need a reusable longer file | Timeline audio concat | 1 to 20 parts, wav default, output up to 1800 s |
Prompt first
The Music Router docs say to steer length in the prompt, and the Music docs call these creative directions rather than guaranteed settings. Treat the requested length as a target and check the artifact. Because a generation costs a fixed $0.125 whatever the length, asking for slightly more than you need and trimming is cheaper than regenerating.
Loop, trim, join
In Timeline 1.0, the render's audio.duration_seconds declares the output length. A soundtrack bed can loop to fill it and fade out for up to 10 seconds. Looping repeats the track, so write a brief that loops well: steady tempo, no big ending, and a short intro. Soft warnings for padded or looped short sources appear on the finished job and are not failures.
To make one longer file from several takes, use timeline audio concat. It joins in the sample domain, with no silence at the seams. Choose wav output, the default, if the file will be joined again or drive lip-sync; mp3 re-adds priming padding at every edge. It costs $0.01 per job.
- A fade-out beats an abrupt loop point for most ads.
- Parts must share a channel layout (
audio_parts_channel_mismatch). - Inputs must already be this workspace's
media.sume.comaudio; import first.
A rule of thumb
These thresholds are my own rule of thumb, not Sume limits. Under 10 percent short: loop and fade. Under half: loop with a calm brief. More than that: ask for a longer track, or join two takes with matching key and tempo so the seam is musical. Always listen to the seam. See Jobs and results for polling each job.
Sources
Related posts
More in Developers
- AI video API: seed and size return 400 on Sume; what to send instead
No Sume video model accepts seed, and size returns 400 unsupported_parameter. Use resolution and aspect_ratio, and keep the prompt and frames to redo a take.
- Vendors swap GPUs; keep one video job shape
Luma said on Jul 23, 2026 it runs video-to-video inference on AMD and Tensorwave. Your client should not care: one Sume job shape covers every model.
- Profit per SKU across channels: add AI media cost from Sume
Amazon lists cross-channel profitability as upcoming. Add what your product images and clips cost per SKU by reading Sume's usage ledger by job id.
- Animate a still by API: first-frame jobs on Sume
fal lists FLUX 3 as animating one still into video. On Sume you send the still as a first frame to a video model; this Python script submits and polls.
Written by Sume