Write the Lyria brief from your shot list: a Python timestamp builder
Lyria responds to [m:ss-m:ss] section markers. A short Python script turns a 6-shot, 59-second shot list into a brief under the 5000-character limit.

Section markers like [0:00-0:30] Intro: ... are the way the Sume Music docs say to steer structure, and the length goes in the prompt because there is no duration field. The script below builds that brief from a shot list, checks that the shots leave no gap, and checks the 5000-character prompt limit, so the music changes where the picture changes.
Google's Lyria page says the model responds well to section tags and to timestamp markers such as [0:00 - 0:10] (read 2026-10-08). It also says generation is single-turn, so a bad brief means a new generation at $0.125, which is why the checks run before the request.
The script
It uses only the standard library and builds the request body. Add your own HTTP call using the Music Router request shape, with an Idempotency-Key header.
import json
SHOTS = [ # (start s, end s, what the music does)
(0, 8, "Intro: sparse Rhodes, no drums"),
(8, 18, "Verse: brushed drums enter, bass follows"),
(18, 27, "Lift: muted trumpet answers the Rhodes"),
(27, 39, "Chorus: full groove"),
(39, 49, "Break: bass and claps only"),
(49, 59, "Outro: groove returns, then a clean tail"),
]
def mmss(t: int) -> str:
return f"{t // 60}:{t % 60:02d}"
assert all(a[1] == b[0] for a, b in zip(SHOTS, SHOTS[1:])), "gap or overlap"
sections = " ".join(f"[{mmss(a)}-{mmss(b)}] {text}." for a, b, text in SHOTS)
brief = f"Warm lo-fi hip hop, 84 BPM, C minor. A 59-second track. {sections} Instrumental, no vocals."
assert len(brief) <= 5000, "Music prompts stop at 5000 characters"
body = {"model": "sume/music-auto", "prompt": brief}
print(len(brief), "characters")
print(json.dumps(body)[:120], "...")What the script guards
Three rules in the docs motivate the checks.
- Prompt length is 1 to 5000 characters. The 6-shot brief above is 357 characters, so a 59-second Short never gets near the limit: at roughly 60 characters a shot, you would need more than 80 shots to reach it.
durationandduration_secondsare rejected, so the 59 seconds goes in the text asA 59-second track, and the last marker ends at 0:59.- Exclusions go in the positive prompt: the brief ends with
Instrumental, no vocals, because a non-emptynegative_promptreturns HTTP 400.
Why the gap check matters
A shot list drawn by hand tends to drift: one shot ends at 18 and the next starts at 19, and the brief then describes a second of music that the picture never shows. The assertion compares each end with the next start, so the brief covers the full length of the Short with no hole and no overlap. A brief with an overlap is the worse case, because two markers can claim the same second and the model has to guess.
The check costs nothing and a mistake costs $0.125 plus the time to listen. Across a 24-video campaign that is $3.00 of generations at risk if every brief is wrong, which is why the builder belongs in the pipeline, not in a person's head.
Variations worth trying
Change the first sentence per market or per product and keep the section lines, so a family of beds shares a structure that fits the same cut. The Music 1.0 page suggests changing the genre family, the tempo (at least 12 BPM apart) and the lead instrument for scenes that should contrast, and keeping continuity when a project wants one score.
If a still frame from the accepted scene exists, pass it as image_url, which the docs allow as one public HTTPS image. It does not change the price, which stays at $0.125 per accepted generation.
Using it with a cut list
Take the shot lengths from your edit. In the example the cuts are at 8, 18, 27, 39, 49 and 59 seconds, and the music changes on each. If your cut list is in frames, convert it to seconds first and round to whole seconds, since m:ss has no fractions.
Treat the markers as intent. The Music 1.0 page says the axes are creative directions, not guaranteed values, and to examine the generated audio. If the break lands at 0:41 instead of 0:39, you can move the video cut by two seconds, or cut the bed with a Timeline audio split. A new generation costs $0.125, a split $0.01.
Sources
Related posts
More in Developers
- 11 common sizes vs GPT Image 2.5 custom rules: 6 pass, 5 fail
GPT Image 2.5 custom sizes need edges in multiples of 16, an edge up to 3840, 655,360 to 8,294,400 pixels and 3:1 or less. 1280x720 passes; 1920x1080 fails.
- A $12.50 balance fails a 15-second Seedance 2.0 1080p job by 26 cents
A 15-second Seedance 2.0 job at 1080p holds $12.7575, so a $12.50 balance gets 402 insufficient_credits; 480p ($2.64) and 720p ($5.67) go through.
- A 1,300-second narration fails tts_duration_exceeded: split the script
Sume TTS fails a request whose audio runs past 1,200 seconds with tts_duration_exceeded and captures no credit. Split the script at sentence ends.
- 15 Ideogram 4.5 edits with 5 images each: $1.125 at medium on Sume
Ideogram 4.5 on Sume edits the first image and takes up to four more as references. Fifteen medium edits cost 15 x $0.075 = $1.125.
Written by Sume