One score or contrasting cues: briefing Lyria 3.5 per scene

For a video with several scenes, decide first: one consistent score or contrasting cues. Sume's music docs give the rule for each, including a 12 BPM gap.

4 min readSume
All posts

Decide before you write the first prompt: does the video want one consistent score, or a different cue per scene? Sume's Music docs give a rule for each. For contrasting scenes, vary the broad genre family, the tempo by at least 12 BPM, and the lead instrument. For one consistent score, keep the brief and carry continuity forward, passing the accepted scene still as image_url when that helps.

The reason to be this deliberate is the model. Google's Gemini API changelog (read 2026-10-02) lists Lyria 3.5 GA on 2026-09-03 with full-length songs, so one request can return a long piece, and the Sume music endpoint takes no seed, so a re-run will not reproduce the first take.

When do I want contrasting cues?

Use them when scenes change mood: a calm product reveal, then a loud comparison, then a quiet call to action. Generate one track per scene and vary three things, so the cues do not blur together.

Per-scene variation rule. Sume Music 1.0 docs, read 2026-10-02.
SettingContrasting scenesOne consistent score
Genre familyChange it between scenesKeep it
TempoAt least 12 BPM apartKeep it
Lead instrumentChange itKeep it
image_urlOptional, the scene's stillPass the accepted scene still when appropriate

How do I keep one score consistent?

Reuse the same genre, tempo, key and instrument line in every request, and change only the arc clause. Because there is no seed, treat each generation as a new take and keep the one you accept. Re-generating the same text will not give you the same track, so store the Sume media.sume.com URL of the accepted file.

How do the tracks get under the video?

Sume's Timeline 1.0 takes an optional soundtrack bed with url, gain_db, loop, fade_out_seconds up to 10, and duck_db from 0 to 20; ducking needs a real audio spine, not silence. That lets the score sit under narration without a manual mix.

What does it cost?

Music 1.0 lists a fixed $0.125 per accepted generation, whatever the prompt length or image input. Three scene cues cost three generations, plus any retakes.

How do I name and store the cues?

Name each file with the scene number and the brief version, and keep the accepted media.sume.com URL in a small table beside the script. When someone asks for the second scene's music to be softer, you can find its brief, change one clause, and regenerate only that cue.

Keep the rejected takes for a while too. A take that was wrong for one scene is sometimes right for another.

How do I decide which approach fits?

Ask what the viewer should feel across the video. If the answer is one steady mood, such as calm trust, use one score. If the answer is a sequence of moods, such as curiosity, tension and relief, use contrasting cues and change tempo, genre family and lead instrument between them.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume