Where are the lyrics in a Lyria 3.5 result? Gemini vs Sume
Google returns Lyria 3.5 lyrics and song structure as text beside the audio. Sume puts model-reported lyrics or a section map in result.lyrics when present.

On Google's Gemini API the Lyria 3.5 model page lists its outputs as audio (MP3) and text (lyrics), and the music guide says you read the song's lyrics and structure from interaction.output_text. On Sume the same information, when the engine provides it, is in result.lyrics on the finished job, next to the audio artifact in result.artifacts[]. The Music Router docs say result.lyrics carries the model-reported lyrics or section map when present.
What does Google say you get back?
From the Lyria 3.5 model page: input is text and image, output is audio (MP3) and text (lyrics), and the input token limit is 131,072. From the music generation guide: output_audio returns the last audio block and output_text returns lyrics and structure, with a note that for interleaved responses you may need to iterate over the steps yourself.
| Item | Gemini API (Interactions) | Sume job result |
|---|---|---|
| Audio | interaction.output_audio (base64 data) | result.artifacts[] with type audio, a media.sume.com URL |
| Lyrics or structure | interaction.output_text | result.lyrics, when present |
| Default format | MP3; WAV by response_format | Typically audio/mpeg |
Can I trust result.lyrics as a timing map?
No. The Music 1.0 docs say provider lyrics may describe tempo and structure but that this is model-reported metadata, not an audio measurement. If you need to cut at a section boundary, find the boundary in the audio. For lyric captions, the stored guide Lyric video captions for an AI song shows the script_text route.
How do I know whether an engine returns lyrics?
Ask the catalog: GET /v1/music-router/models rows include capabilities with text_to_music, image_conditioning, lyrics and max_prompt_characters. Check the lyrics flag for the id you plan to pin; with sume/music-auto, read the field on the result and handle its absence.
Sources
Related posts
More in Models
- Nano Banana negative prompt: describe what you want instead
Google's Gemini image guide says to write semantic negative prompts: describe an empty street, not 'no cars'. Sume's image request has no negative field.
- Nano Banana prompt languages: Korean and Japanese, per Google
Google lists the languages Gemini image models work best in, including ko-KR and ja-JP. What that means for a Korean or Japanese prompt sent through Sume.
- Nano Banana reference limits by model: objects, characters, style
Google splits Nano Banana references by model: Nano Banana 2 takes 10 objects, 4 characters, 3 style images; Pro takes 6, 5, 3. Sume uses one list.
- Nano Banana text in images: write the copy first, then render
Google's tip for text in Gemini images: settle the wording first, then ask for the image. How to do that with one Sume request and a short checklist.
Written by Sume