Score a multi-scene video with AI music: tempo, key, lead instrument

How to brief Sume Music for a video with several scenes: one consistent score or contrasting cues, with tempo, key and lead-instrument rules from the docs.

5 min readSume
All posts

Write one music brief per scene. For contrasting scenes, change the broad genre family, set tempos at least 12 BPM apart and give each scene a different lead instrument; for one consistent score, keep the key, palette and tempo family and vary only the arrangement. This is the guidance in Sume's Music 1.0 docs, and it works with Sume's Music Router, which runs Lyria 3.5 today.

Music takes no seed, temperature, guidance or duration parameter, so the prompt is your only control. The docs call these briefs creative directions, not guaranteed output settings, so listen to every result.

What goes into a scene brief?

The docs name seven axes: emotion, genre or lineage, tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and era or production style. Close with one clause, "Instrumental, no vocals," and add "no spoken word" only when narration will sit on top.

A scene-specific brief beats a mood word because it removes the generic choices the model would otherwise make. The docs' own example reads: hushed and slightly melancholic neo-soul nocturne, 72 BPM, D minor, Rhodes through tape wow, soft sub bass, brushed snare, a single muted trumpet, a sparse first half where the trumpet answers the Rhodes from 0:12, dry and close 1998 production.

How do I make scenes feel different?

When a video moves between a calm opening, a busy product demo and a closing call to action, a single bed flattens all three. The docs' rule for contrasting scenes has three levers, shown with an example plan below.

A three-scene plan following the contrast rule in the Music 1.0 docs, read 2026-10-02
SceneGenre familyTempoKeyLead instrument
Opening, quietChamber folk72 BPMD minorCello
Demo, busySynthwave128 BPMA minorAnalog arpeggio synth
Closing, warmNeo-soul92 BPMF majorRhodes

How do I keep one consistent score?

If the client or the format asks for a single score, preserve continuity instead: repeat the key and mode, reuse the same two or three instruments, keep tempos within a narrow band and change only the arc, such as a sparser first pass and a fuller final one. Sume does not offer a way to pin a motif between generations, so continuity comes from repeating the same words in each brief, and results will still vary.

The docs add one more tool: pass the accepted scene still as image_url when appropriate. Image conditioning is optional, and the price does not change with it. Use public HTTPS image URLs only.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: score-scene-2" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Driving synthwave, 128 BPM, A minor. Analog arpeggio, gated snare, sub bass. Breakdown to bass and claps at 0:10, full return at 0:14. A 20-second track. Instrumental, no vocals.",
    "image_url": "https://example.com/scene-2.png"
  }'

How do I put the cues on the timeline?

Each generation returns one audio file. To join cues, use timeline audio concat on up to 20 Sume-hosted files; it joins in the sample domain with no gaps, at $0.01 per job. Use WAV as the output format when the file will be joined again. To mix under a voiceover, send the joined file as soundtrack in a Timeline 1.0 render with duck_db between 0 and 20 and a fade_out_seconds of up to 10.

Because the length is steered by words, a requested 20-second cue may come back longer or shorter. Read the duration from the result and set each duration or source_in on the concat part so that the seams land where the scene cuts do. Inside a Timeline 1.0 render, audio.parts whose lengths sum to less than duration_seconds are refused with audio_parts_shorter_than_duration.

What does it cost, and what happens on a rejection?

Each accepted generation is a fixed $0.125, so a three-scene score with two takes per scene is about $0.75, plus $0.01 for the join. If a request is rejected for policy, the docs say to revise the flagged content while keeping the musical brief, and to retry only within the budget you were authorised to spend; do not strip the request down to a generic bed. A non-empty negative_prompt is a 400 (negative_prompt_unsupported), so write exclusions into the positive prompt.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume