AI music for a 30-second video: no duration field, so prompt it
Sume's Music Router rejects duration and duration_seconds. Steer length in the prompt, then cut the track to your video with a timeline audio split.

Describe the length in the prompt, generate, and cut the result to fit your video. The Music Router does not accept a duration: duration and duration_seconds are rejected, and the docs tell you to steer length in the prompt itself, with examples like a 2-minute track or timed section labels such as [0:00-0:30] Intro.
For a 30-second video that means asking for a short track with a clear ending, then trimming whatever you get back. The cut uses Sume's timeline audio split.
How do you write the length into the prompt?
The prompt takes 1 to 5000 characters. Put exclusions in the positive text too, because negative_prompt is not supported, as covered in the negative prompt post. A prompt for a short spot could say: a 30-second upbeat acoustic track, bright intro in the first 5 seconds, clean ending at 30 seconds, no vocals.
| Control | Supported | Note |
|---|---|---|
duration or duration_seconds | No, rejected | Do not send them |
| Length wording in the prompt | Yes | For example a 2-minute track |
| Timed section labels in the prompt | Yes | For example [0:00-0:30] Intro: ... |
negative_prompt | No | State exclusions in the positive prompt |
| Prompt length | 1 to 5000 characters | Per the docs |
What if the track is longer than the video?
Prompt wording steers length but does not guarantee it, so plan for a track that runs over. Timeline audio with operation: split slices one audio file into ranges, each with its own audio_url. The source must be audio on your workspace's media.sume.com, so import it first. A job is priced at $0.01 flat under the current estimate, and the docs say to confirm in the catalogue.
A request is a url and ranges, for example one range from 0 to 30 seconds. Keep the default wav output if the file will be joined again, and use mp3 only for a final file, since the docs note mp3 re-adds padding at every edge.
What about the ending?
Music 1.0 is retiring gradually, per the Music Router page, so new work should use the router. Read the Music Router docs for the routable model ids and the request body.
- A hard cut at 30 seconds can land mid-phrase. Ask for a clean ending in the prompt so the cut falls near a natural stop.
- If it still sounds abrupt, generate again with the section labels moved earlier.
- Use the same split job to take a second range, such as a 3-second tail, if you want to build a fade yourself in your editor.
Does it help to generate more than one take?
It does, because length steering through a prompt is not exact. Generate two or three takes with slightly different wording, and pick the one whose natural ending lands closest to your video length. The split step is cheap, so choosing a track that needs only a small trim is better than forcing a long one to fit.
Name the tracks with the prompt text and the date, so you can find the best one again for a later cut.
Sources
Related posts
More in Media tools
- Amazon audio ads, 3 MB cap: 30 s mono WAV fits, stereo does not
Amazon audio ads allow 10 to 30 seconds and 3 MB. In 16-bit PCM, 30 s mono at 44.1 kHz is 2.65 MB and stereo is 5.29 MB. Detach it with Sume audio-detach.
- Burn feature text onto silent product clips with caption cues
A silent product clip fails speech captions. Pass your own cues (text, start, end) to Sume's video-captions endpoint and burn feature text on for TikTok Shop.
- Caption a silent AI video: fixing caption_no_speech
A silent clip fails POST /v1/video-captions with caption_no_speech. Send cues with text, start and end to burn authored captions without speech-to-text.
- Caption a silent Seedance 2.5 clip with authored cues
A clip with no speech fails Sume caption jobs as caption_no_speech. Pass cues with text, start and end seconds instead; $0.20 per job up to 60 seconds.
Written by Sume