Explainer music that changes per scene: three Lyria briefs, one concat
Make each scene's music a separate Lyria 3.5 brief at least 12 BPM apart, then join them gaplessly with Timeline audio concat for $0.01.

For an explainer where each scene needs different music, make one Lyria 3.5 track per scene with POST /v1/music-router/generate, then join them with operation: "concat" on POST /v1/timeline-1.0/audio. The concat is gapless, sample-domain and costs $0.01 flat. Steer each track's length in its prompt, because duration and duration_seconds are rejected.
Source: Music Router and Timeline audio, read 2026-10-04.
Three briefs
Sume's music brief template says contrasting scenes should vary genre family, tempo by at least 12 BPM, and lead instrument. Three briefs below follow that.
| Scene | Prompt summary | BPM | Lead |
|---|---|---|---|
| Problem | Tense minimal electronic, a 30-second track | 78 | Pulsing synth bass |
| Solution | Bright acoustic pop, a 30-second track | 102 | Plucked guitar |
| Call to action | Confident funk, a 30-second track | 118 | Brass stabs |
Submit the first brief
The BPM gaps are 24 and 16, both above the 12 BPM rule. Put exclusions in the positive prompt, since a non-empty negative_prompt is rejected. Each request is one job, so use an Idempotency-Key per scene.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: explainer-music-problem-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Tense minimal electronic, 78 BPM, A minor. Pulsing synth bass, sparse clicks, no drums until 0:20. A 30-second track. Instrumental, no vocals."
}\Join them
When all three tracks are ready, read each audio artifact from result.artifacts[]. Tracks must be on media.sume.com and share a channel layout, or concat fails with audio_parts_channel_mismatch. Concat takes 1 to 20 parts.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: explainer-music-concat-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/problem.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/solution.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/cta.wav" }
]
}\Line up the picture
Lyria does not guarantee exactly 30 seconds, so read segments[] from the concat result. It gives each part's start and duration_seconds, which are the offsets to re-base Timeline 1.0 video[].start against. Use the merged file as audio.url on the render.
Sources
Related posts
More in Use cases
- Music clip with Seedance 2.5: audio reference plus performer image
A Seedance 2.5 request on Sume takes reference_audio_urls only with a reference image or video. Build a 15 to 30 s music-driven clip and see the cost.
- Muted product loop in four languages: burn captions from cues
A silent clip fails speech-to-captions as caption_no_speech. Pass cues with text, start and end to burn four translations at $0.20 a job.
- Naver Clip video: a 9:16 Sume clip with Korean captions
Making a vertical 9:16 clip for Naver Clip? Generate it on Sume, then burn Korean captions with the korean-ad style and language ko.
- Near-duplicate AI images: perceptual-hash dedupe before human review
Four-up image batches often contain near twins. Use a 64-bit difference hash in Pillow to drop duplicates before a person reviews them. Python for Sume results.
Written by Sume