Write the music brief after the voiceover, using TTS timestamps

Lyria on Sume takes no duration field, only time ranges in the prompt. Generate the voiceover first, read its word timings, then write the section markers.

5 min readSume
All posts

Make the voiceover first, then write the music brief from its timings. Sume's music endpoint has no duration or duration_seconds field, so length and structure live in the prompt as time ranges like [0:00-0:08] Intro: .... A TTS job that returns timestamps.words and sentence segments tells you exactly where those ranges should fall.

Why the order matters

If you write the music first, you guess the voice's length and then fit speech to the track. Speech is the harder constraint: it has words that must land on screen events. Generating it first turns a guess into measured times, and Sume's TTS can tell you where each sentence starts and ends.

Lyria 3.5 lists a $0.08 per song price on Google's page (read 2026-10-07), and Sume charges a fixed $0.125 for each accepted generation. Either way, a track written against measured timings avoids a paid retake caused by a bad length guess.

Read the timings

Request timestamps.words: true and segmentation.mode: sentence on the TTS job. Segments come back gapless, each with a start and end in seconds. Suppose a 24-second read has four sentences.

Example: TTS segment times mapped to music sections - illustrative values, format per Sume docs (read 2026-10-07)
SegmentStartsEndsMusic section marker
1 Hook0.0 s5.5 s[0:00-0:05] Intro: sparse Rhodes, no drums
2 Problem5.5 s12.0 s[0:05-0:12] Build: soft kick enters
3 Offer12.0 s19.0 s[0:12-0:19] Lift: bass and claps, full band
4 Call to action19.0 s24.0 s[0:19-0:24] Resolve: strip to Rhodes, end on tonic

Assemble the prompt

Music 1.0 documents the seven-axis brief: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, an arc with a named moment, and an era or production note. End with "Instrumental, no vocals," and add "no spoken word" because you are placing it under narration. Round section edges to whole seconds, since the docs say the controls are creative directions, not guaranteed output values.

[0:00-0:05] Intro: sparse Rhodes through tape wow, no drums.
[0:05-0:12] Build: soft kick and brushed snare enter.
[0:12-0:19] Lift: warm bass and claps, full band.
[0:19-0:24] Resolve: strip back to Rhodes, end on the tonic.
Hushed, hopeful neo-soul, 88 BPM, D major. A 24-second track. Instrumental, no vocals, no spoken word.

Check the result, then mix

  • Listen for the lift near 0:12. If it lands at 0:15, adjust the brief and take one more generation, knowing there is no seed.
  • Place the voice on the timeline over the music with the voice louder; music ducking is a mix decision, not a music parameter.
  • Send an image_url of the accepted scene still only if the visual mood should steer the track.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume