Lyria prompt section markers: [0:00-0:30] Intro on Sume Music
Sume Music has no duration field. You shape a track with section markers like [0:00-0:30] Intro and a stated length in the prompt itself.

Sume Music takes its structure from the prompt. There is no duration or duration_seconds field; the docs say the API rejects both. To get a structured track, write the length in words ("a 2-minute track") and add section markers in the form [0:00-0:30] Intro: .... Music 1.0 runs on Google Lyria 3.5 and makes full-length structured songs of up to a few minutes, so the markers are how you divide that time.
A marked-up brief
The marker syntax comes straight from the Sume docs. The musical axes below are the docs' seven-axis brief: emotion, genre, tempo as a number, key and mode, two to four textured instruments, an arc with one named moment, and era or production. They are creative directions, not guaranteed output values, so listen to the result.
A 1-minute track. 96 BPM, A minor, warm analog synthwave.
[0:00-0:15] Intro: soft pad and a muted bass pulse, sparse.
[0:15-0:40] Verse: add gated drums and a plucked lead, restless.
[0:40-1:00] Outro: strip back to pad, slow fade.
Dry and close, 1986 production. Instrumental, no vocals.What the API does with it
Nothing special. The prompt is sent as text (1 to 5000 characters), and the accepted request is charged a fixed price. The markers never change the price, which stays the same for a 30-second brief and a 3-minute one.
| Item | Value |
|---|---|
| Prompt length | 1 to 5000 characters |
duration / duration_seconds | Rejected |
Non-empty negative_prompt | HTTP 400 negative_prompt_unsupported |
| Price per accepted generation | $0.125 flat |
| Engine today | Google Lyria 3.5 (routed via sume/music-auto) |
Why lyrics output matters
When the provider returns them, result.lyrics carries model-reported lyrics or a section map. The docs call this metadata, not an audio measurement. Do not use it to prove the track matches your markers; compare by ear, or measure the audio artifact yourself.
Retry rules
If a policy rejection comes back, change the flagged content but keep the musical brief, and retry only within the budget you were given. Do not shrink the request to a generic bed, because the brief is the part you pay for. Use an Idempotency-Key so a network retry does not create a second $0.125 job.
A checklist before you submit
Check three things. First, the marker times should add up to the length you wrote in the sentence at the top of the prompt, so a "1-minute track" ends its last marker at 1:00. Second, keep each section to one clear change, since the model reads the line as direction, not as a score. Third, end with the instrumental clause so vocals do not appear in a bed.
- Wrap the request in an
Idempotency-Key. - Poll
GET /v1/jobs/:id/status, then readresult.artifacts[]for theaudioentry. - Listen to the result and compare the section changes by ear.
Sources
Related posts
More in Media tools
- MAI-Voice 24 kHz 160 kbps MP3 vs Sume TTS 44.1 kHz 128 kbps: mixing
Microsoft's MAI-Voice example saves 24 kHz 160 kbps mono MP3; Sume TTS defaults to 44.1 kHz 128 kbps MP3 or WAV. What the numbers mean for mixing and file size.
- Meta feed sound is optional but recommended: add a ducked music bed
Meta says captions and sound are optional but recommended on Facebook Feed video. Add a music bed under a voice with Sume Timeline soundtrack and duck_db.
- MiniMax H3 on Sume: 9 reference images, first 5 free, 4 extra $0.08
Sume lists MiniMax H3 with up to 9 reference images; the first 5 are free and each extra is $0.08 list, so 4 extra is $0.32 list. Vidu Q4 lists 15.
- Move burned captions: design.placement.anchor_ratio on Sume
Sume caption styles have default heights. Send design.placement.anchor_ratio to move the line, as a fraction of frame height. Which styles accept it.
Written by Sume