Write the music brief after the voiceover, using TTS timestamps
Lyria on Sume takes no duration field, only time ranges in the prompt. Generate the voiceover first, read its word timings, then write the section markers.

Make the voiceover first, then write the music brief from its timings. Sume's music endpoint has no duration or duration_seconds field, so length and structure live in the prompt as time ranges like [0:00-0:08] Intro: .... A TTS job that returns timestamps.words and sentence segments tells you exactly where those ranges should fall.
Why the order matters
If you write the music first, you guess the voice's length and then fit speech to the track. Speech is the harder constraint: it has words that must land on screen events. Generating it first turns a guess into measured times, and Sume's TTS can tell you where each sentence starts and ends.
Lyria 3.5 lists a $0.08 per song price on Google's page (read 2026-10-07), and Sume charges a fixed $0.125 for each accepted generation. Either way, a track written against measured timings avoids a paid retake caused by a bad length guess.
Read the timings
Request timestamps.words: true and segmentation.mode: sentence on the TTS job. Segments come back gapless, each with a start and end in seconds. Suppose a 24-second read has four sentences.
| Segment | Starts | Ends | Music section marker |
|---|---|---|---|
| 1 Hook | 0.0 s | 5.5 s | [0:00-0:05] Intro: sparse Rhodes, no drums |
| 2 Problem | 5.5 s | 12.0 s | [0:05-0:12] Build: soft kick enters |
| 3 Offer | 12.0 s | 19.0 s | [0:12-0:19] Lift: bass and claps, full band |
| 4 Call to action | 19.0 s | 24.0 s | [0:19-0:24] Resolve: strip to Rhodes, end on tonic |
Assemble the prompt
Music 1.0 documents the seven-axis brief: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, an arc with a named moment, and an era or production note. End with "Instrumental, no vocals," and add "no spoken word" because you are placing it under narration. Round section edges to whole seconds, since the docs say the controls are creative directions, not guaranteed output values.
[0:00-0:05] Intro: sparse Rhodes through tape wow, no drums.
[0:05-0:12] Build: soft kick and brushed snare enter.
[0:12-0:19] Lift: warm bass and claps, full band.
[0:19-0:24] Resolve: strip back to Rhodes, end on the tonic.
Hushed, hopeful neo-soul, 88 BPM, D major. A 24-second track. Instrumental, no vocals, no spoken word.Check the result, then mix
- Listen for the lift near 0:12. If it lands at 0:15, adjust the brief and take one more generation, knowing there is no seed.
- Place the voice on the timeline over the music with the voice louder; music ducking is a mix decision, not a music parameter.
- Send an
image_urlof the accepted scene still only if the visual mood should steer the track.
Sources
Related posts
More in Use cases
- Yoga studio class promo on Wan 3.0: 38 cents to $1.50 for 6 seconds
A 6-second yoga class promo on Sume's wan-3.0 costs 38 cents at 480p, 75 cents at 720p and $1.50 at 1080p. Pick the tier by where the clip will play.
- YouTube bumper ad 5-6 seconds: generate at 6 or trim on Sume
YouTube's help page puts bumper ads at 5 to 6 seconds. Which Sume video models can generate a 6 second clip directly, and when a $0.02 trim is the better route.
- YouTube Shorts ad under 60 seconds: stitch clips with Timeline
YouTube suggests Shorts ads under 60 seconds. Sume Timeline 1.0 stitches clips into one MP4 at $0.10 per output minute, so 60 seconds costs $0.10.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
Written by Sume