Sync music to avatar scenes: hook, demo, CTA timestamps in the prompt

Match a Lyria bed to a 12-second multi-scene avatar video by writing [0:00-0:03] section markers into the Music Router prompt. A worked example and cost.

5 min readSume
All posts

Write the avatar video's scene durations into the music prompt as section markers. A 12-second video with a 3-second hook, a 4-second silent demo and a 5-second call to action becomes [0:00-0:03] Hook, [0:03-0:07] Demo and [0:07-0:12] CTA in the Music Router prompt. Sume's docs support this style: they say to steer length and structure in the prompt with markers like [0:00-0:30] Intro: ..., since there is no duration parameter.

The cost is $0.125 for the track (read 2026-10-03). The avatar video is billed separately per second, so a 12-second Plus clip is $2.94.

Start from the video plan

The avatar video docs describe multi-scene video_inputs: each scene has an id, a voice (spoken text with a duration, or silence with a required duration) and a background. Total planned duration must land in 4 to 60 seconds. Those scene durations are your music timeline: they are known before you render anything.

Write the table first. For each scene, note the job of the music: build tension under the hook, drop to nearly nothing for the demo so the viewer hears the product, lift for the call to action.

12-second plan, read 2026-10-03
SceneMarkerDurationMusic job
hook0:00-0:033 sPickup, one rising motif
demo0:03-0:074 sSparse, drums out, bass only
cta0:07-0:125 sFull return, resolve on the last beat

The prompt

Combine the brief axes (tempo, key, instruments, arc) with the section markers. Treat the markers as creative direction, not guaranteed timing: the Music docs say these are not output settings, and that you should verify the generated audio. If a section lands a second late, trim or offset the track in Timeline rather than regenerating immediately.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-12s-bed-001" \
  -d '{
    "prompt": "Confident synth-pop, 108 BPM, E minor. [0:00-0:03] Hook: rising arpeggio, tight kick. [0:03-0:07] Demo: bass and soft pad only, no drums. [0:07-0:12] CTA: full band returns, final chord on the last beat. A 12-second track. Instrumental, no vocals."
  }'

Checking the result

Because a regeneration costs so little, the practical method is to generate two or three candidates and keep the best.

  • Listen for the drop at 0:03: the demo section should be quiet enough for narration or a screen recording.
  • Check the ending: a bed that fades out instead of resolving feels unfinished at 12 seconds.
  • Read result.lyrics, which may carry a model-reported section map, but remember it is metadata, not an audio measurement.
  • If it is close, trim in Timeline; if the structure is wrong, change the markers and regenerate for another $0.125.

Silence beats and music

A silence scene in video_inputs is a non-speaking beat with a required duration, and it is the perfect place for the music to carry the moment. In the example the 4-second demo scene is silent, so the music is the only audio. Make the section feel intentional: a clean drum-out and a bass line, not an accidental gap. If you instead want narration over the demo, change the scene to a spoken one and tell the music prompt to keep its mids sparse.

Also remember that avatar scene durations are planned, not guaranteed, estimates. If the rendered video comes out a second longer or shorter than the plan, the music markers will drift. Check the final duration with the video inspection route or your player before mixing, and trim the track rather than stretching it.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume