Sync music to avatar scenes: hook, demo, CTA timestamps in the prompt
Match a Lyria bed to a 12-second multi-scene avatar video by writing [0:00-0:03] section markers into the Music Router prompt. A worked example and cost.
Write the avatar video's scene durations into the music prompt as section markers. A 12-second video with a 3-second hook, a 4-second silent demo and a 5-second call to action becomes [0:00-0:03] Hook, [0:03-0:07] Demo and [0:07-0:12] CTA in the Music Router prompt. Sume's docs support this style: they say to steer length and structure in the prompt with markers like [0:00-0:30] Intro: ..., since there is no duration parameter.
The cost is $0.125 for the track (read 2026-10-03). The avatar video is billed separately per second, so a 12-second Plus clip is $2.94.
Start from the video plan
The avatar video docs describe multi-scene video_inputs: each scene has an id, a voice (spoken text with a duration, or silence with a required duration) and a background. Total planned duration must land in 4 to 60 seconds. Those scene durations are your music timeline: they are known before you render anything.
Write the table first. For each scene, note the job of the music: build tension under the hook, drop to nearly nothing for the demo so the viewer hears the product, lift for the call to action.
| Scene | Marker | Duration | Music job |
|---|---|---|---|
| hook | 0:00-0:03 | 3 s | Pickup, one rising motif |
| demo | 0:03-0:07 | 4 s | Sparse, drums out, bass only |
| cta | 0:07-0:12 | 5 s | Full return, resolve on the last beat |
The prompt
Combine the brief axes (tempo, key, instruments, arc) with the section markers. Treat the markers as creative direction, not guaranteed timing: the Music docs say these are not output settings, and that you should verify the generated audio. If a section lands a second late, trim or offset the track in Timeline rather than regenerating immediately.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-12s-bed-001" \
-d '{
"prompt": "Confident synth-pop, 108 BPM, E minor. [0:00-0:03] Hook: rising arpeggio, tight kick. [0:03-0:07] Demo: bass and soft pad only, no drums. [0:07-0:12] CTA: full band returns, final chord on the last beat. A 12-second track. Instrumental, no vocals."
}'Checking the result
Because a regeneration costs so little, the practical method is to generate two or three candidates and keep the best.
- Listen for the drop at 0:03: the demo section should be quiet enough for narration or a screen recording.
- Check the ending: a bed that fades out instead of resolving feels unfinished at 12 seconds.
- Read
result.lyrics, which may carry a model-reported section map, but remember it is metadata, not an audio measurement. - If it is close, trim in Timeline; if the structure is wrong, change the markers and regenerate for another $0.125.
Silence beats and music
A silence scene in video_inputs is a non-speaking beat with a required duration, and it is the perfect place for the music to carry the moment. In the example the 4-second demo scene is silent, so the music is the only audio. Make the section feel intentional: a clean drum-out and a bass line, not an accidental gap. If you instead want narration over the demo, change the scene to a spoken one and tell the music prompt to keep its mids sparse.
Also remember that avatar scene durations are planned, not guaranteed, estimates. If the rendered video comes out a second longer or shorter than the plan, the music markers will drift. Check the final duration with the video inspection route or your player before mixing, and trim the track rather than stretching it.
Sources
Related posts
More in Use cases
- Taboola Realize video ad specs: 15 s motion ad vs 90 s video
Realize (Taboola) lists a 15-second motion ad and a video spec of 6 to 30 s, 90 s max. The sizes, files and how to cut a product clip for each with Sume.
- Tag products in a YouTube Short: the sound rule and a Sume clip
Product tags in a Short: tag in the order products appear, use a Shopping sound or no sound, and block reasons like copyright claims. Plan the clip with Sume.
- Teams translated captions vanish after the meeting: keep them
Microsoft says Teams translated captions and transcripts are only available during the meeting. Caption the recording afterwards with Sume STT and burned cues.
- Template bulk edits on YouTube Shorts: what to vary per row
YouTube's Oct 1, 2026 originality update names template-based bulk changes as not original. How to make each row of a Sume bulk run differ in substance.
Written by Sume