AI song structure tags: [Verse], [Chorus] and timestamps
How to structure an AI song prompt for Lyria 3.5 with section tags and timestamps, and how the same prompt goes through Sume's Music Router.

Structure an AI song by putting section tags or timestamps in the prompt. Google's Lyria page says to use tags like [Verse], [Chorus] and [Bridge], and also documents timestamp sections such as [0:00 - 0:10] Intro: .... Sume's docs show the same timestamp style, [0:00-0:30] Intro: ..., and send the text to Lyria 3.5 through the Music Router.
Google's facts are from its Lyria page; Sume's from Music 1.0 and the Music Router. Read 2026-09-29.
Which structure syntax works?
| Syntax | Use | Source |
|---|---|---|
[Verse], [Chorus], [Bridge] | Guide the song form; place lyrics under each tag | |
[0:00 - 0:10] Intro: ... | Time-stamped sections | |
[0:00-0:30] Intro: ... | Steer length and form | Sume docs |
| "a 2-minute track" | Steer length in words | Sume docs |
What does a structured request look like?
The tags are plain text inside prompt; there is no separate structure field. Combine them with tempo, key and instruments.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: structure-001" \
-d '{
"model": "lyria-3.5",
"prompt": "Anthemic pop rock, 124 BPM, E major, no vocals. [0:00-0:15] Intro: clean guitar. [0:15-0:40] Verse: bass and drums enter. [0:40-1:00] Chorus: full band, big drums. A 1-minute track."
}'Do the sections have to land exactly?
No guarantee. Sume's docs describe prompt directions as creative guidance and tell you to verify the audio; Google notes results can vary between calls. Listen for where the chorus actually starts, and adjust the times in a new take.
Where do I limit the prompt?
A prompt is 1 to 5,000 characters. Structure a song in a few tagged lines rather than a long essay, and put exclusions in the positive text, since a non-empty negative_prompt returns HTTP 400.
Sources
Related posts
More in Media tools
- AI sound effects for video: what Sume lists and what it does not
Sume does not list a dedicated sound-effects model. Sound for video comes from video models with generate_audio, the music route, or your files on the Timeline.
- Amazon online video ad specs: OLV size, length, bitrate
Amazon online video (OLV) ads run 6–120 s in 16:9, at least 1920×1080 and 4 Mbps, with 192 kbps AAC on 2+ channels and up to 500 MB site-served.
- Audio ad specs: Spotify, Amazon, SiriusXM, and YouTube
Audio ad specs by seller: Spotify wants 192–320 kbps at -16 LUFS, Amazon a 10–30 s file up to 3 MB, SiriusXM a 44.1 kHz MP3, YouTube a video.
- Extract 16 kHz mono audio from a video for speech-to-text
Set sample_rate 16000 and channels mono on Sume's audio detach to get the speech-to-text shape from a video. Options, the 900 second cap and the price.
Written by Sume