A four-scene score on Sume Music Router: tempos 24 BPM apart, $0.50
Four 15-second scenes need four contrasting Lyria tracks. The docs ask for at least 12 BPM between scenes. Four prompts at 72, 96, 120 and 144 BPM cost $0.50.

For a four-scene film that should feel like four scenes, send four separate Music Router jobs, change the genre family, the tempo and the lead instrument between them, and keep the tempos at least 12 BPM apart as the Music docs advise. Tempos of 72, 96, 120 and 144 BPM are 24 apart. Four jobs at $0.125 each cost $0.50, whatever the prompt length.
The rule the docs give
The Music 1.0 docs say that for scenes in one project that contrast, you change the broad genre family, the tempo (at least 12 BPM apart), and the lead instrument, and that if the user wants one consistent score you keep continuity instead. They also say the model has no duration parameter, so length is written into the prompt, and the axes in a brief are creative directions, not guaranteed output values.
| Scene | Genre family | Tempo | Lead instrument | Prompt ends with |
|---|---|---|---|---|
| 1 Arrival | chamber folk | 72 BPM | fingerpicked guitar | Instrumental, no vocals. |
| 2 Market | bossa nova | 96 BPM | nylon guitar and brushes | Instrumental, no vocals. |
| 3 Chase | drill | 120 BPM half-time pulse | 808 with long glide | Instrumental, no vocals. |
| 4 Dawn | synthwave | 144 BPM | analog arpeggio | Instrumental, no vocals. |
The request
Send each brief to POST /v1/music-router/generate with model omitted or set to sume/music-auto. Add the length to the prompt text, such as "A 15-second track." Do not send duration or duration_seconds; Sume rejects both. Each scene needs its own Idempotency-Key so a retry does not buy a second track.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: score-scene-3" \
-d '{
"model": "sume/music-auto",
"prompt": "Drill, 120 BPM half-time pulse, E phrygian. 808 with long glide, muted hi-hats, a detuned bell line. A 15-second track. Tight and dry, 2024 production. Instrumental, no vocals."
}'When one consistent score is better
Contrast is a choice, not a rule for every film. If the four scenes are one story and the user wants a single mood, generate one longer track instead, and keep the genre, tempo and lead instrument the same across any follow-ups. The docs suggest passing the accepted scene still as image_url when continuity matters, so the music is conditioned on the same picture. Image URLs must be public HTTPS. The price does not change with the image: it is still $0.125 per accepted generation.
After a rejection
If a prompt is rejected by policy, the docs say to change the flagged content but keep the musical brief, and to retry only within your authorized budget. Do not reduce the request to a generic bed. The catalog says failed or pre-generation-canceled jobs are refunded, but a retry that is accepted is a new generation at $0.125, so a $0.50 score becomes $1.00 if every scene is retried once.
Check the result before you cut
Because length is steered by words, the returned file will not always be 15 seconds. Read the audio artifact from result.artifacts[] and measure it. A track that runs long can be cut with a Timeline audio split range, which is $0.01 per job, and job.request.routed_model tells you which engine ran. At $0.125 a track, a second pass on the one scene that missed costs less than rewriting the whole brief.
Sources
Related posts
More in Media tools
- One frame at 12.5 s: video-frames vs video-inspect
Video frames returns a frame at its source size from clips up to 300 s. Video inspect defaults to 768 px, caps at 2160, and takes clips up to 1,800 s.
- Poster frame at 2.5 seconds: video-frames returns source-size stills
Sume POST /v1/video-frames extracts up to 24 exact stills at source size. Always async (202); a failed instant gives a null url and the job still succeeds.
- Google Ads Shorts: horizontal video serves with blurred edges
Google's Shorts ad specs say horizontal assets serve with blurred edges and only the first 60 seconds play in feed. Render your own 9:16 blur fit instead.
- H3 Max Lip Sync on an 11.2 s line: $0.75, $1.20 or $2.40
An 11.2-second voice line bills 12 seconds on Sume's MiniMax H3 Max Lip Sync: $0.75 at 480p, $1.20 at 768p, $2.40 at 1080p. See the audio window and limits.
Written by Sume