Lyria 3.5 prompts: [Verse] tags or [0:00-0:30] time ranges?
Google's Lyria 3.5 docs show [Verse], [Chorus] and [Bridge] tags. Sume's music docs show time-range markers. What each says, and how to test both on one brief.

Google's Lyria 3.5 docs name section tags in the prompt: [Verse], [Chorus] and [Bridge]. Sume's music docs show a different device for the same model family: time-range markers such as [0:00-0:30] Intro: .... Sume's docs do not describe the Google tags, so treat them as something to test rather than assume.
Both are plain text inside one prompt, so a side-by-side test costs two small jobs.
What each page says
Only statements from the two pages as read on 2026-10-03.
| Topic | Gemini API docs | Sume docs |
|---|---|---|
| Section tags | [Verse], [Chorus], [Bridge] | Not described |
| Time ranges | Not described | [0:00-0:30] Intro: ... markers |
| Length | A couple of minutes, controllable by prompt | Say "a 2-minute track" in the prompt; no duration field |
| Turns | Single-turn only | One prompt per job |
| Image input | Text plus up to 10 images | One optional image_url |
| Prompt size | Not stated in the entry reviewed | 1 to 5000 characters |
Two versions of one brief
The brief below is for a 45-second product video bed. Version A uses the tags from Google's docs. Version B uses the time ranges from Sume's. Both keep the same instruments, tempo and closing clause so the structure device is the only variable.
A (section tags)
Warm lo-fi hip hop, 84 BPM, C minor. Dusty Rhodes, brushed drums, muted trumpet.
[Verse] Sparse Rhodes and brushes, no trumpet.
[Chorus] Trumpet answers the Rhodes, bass thickens.
[Bridge] Drums drop out, Rhodes alone.
Instrumental, no vocals.
B (time ranges)
Warm lo-fi hip hop, 84 BPM, C minor. Dusty Rhodes, brushed drums, muted trumpet. A 45-second track.
[0:00-0:15] Intro: sparse Rhodes and brushes, no trumpet.
[0:15-0:35] Main: trumpet answers the Rhodes, bass thickens.
[0:35-0:45] Outro: drums drop out, Rhodes alone.
Instrumental, no vocals.How to compare them fairly
Submit each version through POST /v1/music-router/generate and listen with the brief in front of you.
- Run each version at least twice. Google's docs say the model is single-turn, so there is no edit step to correct a miss; a miss means a new job.
- Check the section boundaries by ear against the markers. Sume's music docs say a brief is a creative direction, not a guaranteed output setting, so verify the audio.
- Read
result.lyricson the finished job if present. The docs call it model-reported metadata describing structure, not an audio measurement. - Keep the closing clause identical. Sume's docs put exclusions in the positive prompt, for example "Instrumental, no vocals.", because a non-empty
negative_promptreturns a 400.
Which to default to
On Sume, default to the time-range form, since it is the one Sume's docs show and it doubles as a length statement. If you also render the same brief directly against Google's API, keep the tag form for that path and keep both versions in your prompt library under one brief id. The goal is a result you can reproduce from a stored prompt, not a rule about brackets.
Sources
Related posts
More in Media tools
- Seedance 2.5 secondary edit vs trim-and-regenerate on Sume
BytePlus describes timestamp-level edits to Seedance 2.5 clips. Sume has no such edit field, so here is the trim-and-regenerate route and its limits.
- Stability AI's Series B and label backers: what it means for audio
Stability released Stable Audio 3.0 on 5/20/26 and raised a Series B on 8/25/26 with EA, Sony, UMG and WMG. A dated timeline and what to verify for video work.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
Written by Sume