AI music generator with vocals: Lyria 3.5 lyrics by API
Lyria 3.5 can sing. How to ask for vocals or an instrumental, steer lyrics with section tags, and read the lyrics back from a Sume music job.

An AI music generator with vocals writes a sung track from a text prompt, and Lyria 3.5 is one that does. Google says it makes songs with vocals and timed lyrics, generates lyrics in the language of your prompt, and takes section tags like [Verse] and [Chorus]. On Sume, send the prompt to the Music Router and read the lyrics from result.lyrics.
Google's facts are from its Lyria page and Flow Music post; Sume's from the Music Router and Music 1.0 docs. Read 2026-09-29.
How do I ask for vocals or no vocals?
Say it in the prompt. Sume's Music 1.0 docs close briefs with "Instrumental, no vocals." when the track should be a bed, and add "no spoken word" only under narration. For a sung track, describe the voice and the words instead.
- Sung: "Warm female vocal, breathy, close-mic'd. Lyrics about a night train home."
- Instrumental: end the prompt with "Instrumental, no vocals."
- There is no separate vocals switch in the request; the prompt is the control.
Can I write my own lyrics?
Google's page says to use section tags such as [Verse], [Chorus] and [Bridge] to guide the composition, so a prompt can carry your lyrics under those tags. Sume's prompt field takes 1-5,000 characters and the catalog lists lyrics as a capability of every Music Router model. Listen to the result: the docs call prompt directions guidance, not guaranteed settings.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: vocals-001" \
-d '{
"model": "lyria-3.5",
"prompt": "Indie folk, 96 BPM, A major, warm female vocal. [Verse] The last train hums along the bay [Chorus] Carry me home before the day. A 1-minute track."
}'How do I read the lyrics back?
When the job is done, GET /v1/jobs/{id}/result has the audio in result.artifacts[] and, when present, result.lyrics. The docs call this the model-reported lyrics or section map: it is metadata from the model, not a transcript of the audio. To check what is actually sung, run the audio through speech-to-text.
What will not work?
| Limit | Source |
|---|---|
| Requests for a specific artist's voice or copyrighted lyrics are blocked by safety filters | |
| Single-turn: you cannot refine a result with follow-up prompts | |
| Results can differ between calls, even with the same prompt | |
Non-empty negative_prompt returns HTTP 400 | Sume |
No duration field; steer length in the prompt | Sume |
Sources
Related posts
More in Models
- Lyria 3.5 release date: where it is available
Google announced Lyria 3.5 in the Gemini app on 2026-09-04. Where Google says you can use it, and where Sume lists it.
- How long is a Lyria 3.5 song, and how do I set the length?
Lyria 3.5 makes songs of a couple of minutes, and length is set in the prompt. What Google says, what Sume rejects, and prompts that steer duration.
- 2K AI video generator API: MiniMax H3 at 2K on Sume
MiniMax H3 makes video up to 2K. On Sume you ask for resolution 2K on minimax-h3, an upscale of native 768p. Request, price for 10 seconds, and limits.
- MiniMax H3 Max 1080p: a latent refinement from native 768p
MiniMax H3 Max on Sume offers 480p, 768p and 1080p. The 1080p tier is a latent refinement from native 768p. What that means for price and output checks.
Written by Sume