Spanish song lyrics from an AI music prompt: write it in Spanish
Google says Lyria lyrics follow the prompt language. For a Spanish track on Sume, write the prompt and lyrics in Spanish, then read the lyrics section map.

Write the whole prompt in Spanish, including any lyrics you supply, because Google's Lyria documentation says the lyrics follow the language of the prompt. On Sume, send that prompt to the Music Router with lyria-3.5 or sume/music-auto, then check the model-reported section map in the result.
There is no language switch to flip. The language is a property of the text you send, so a Spanish brief with English genre words can drift, and a fully Spanish brief is the safest way to ask for Spanish vocals.
What Google documents
The Gemini API music generation page says Lyria 3.5 supports custom lyrics with section tags such as [Verse] and [Chorus], and that generated lyrics follow the prompt language. It also states that requests for an artist's voice and for copyrighted lyrics are blocked, and that the interface is single-turn only. Output is MP3 or WAV at 44.1 kHz stereo and carries a SynthID watermark.
Single-turn means you cannot reply with a correction to the same session. To change a line, edit the prompt and generate again.
A prompt shape that works
Keep section tags in the same bracketed form Google documents, and keep the instructions in Spanish too, so the model does not see mixed signals about language. A skeleton looks like this in plain text: a short description of style and mood, followed by [Verse] with two or four lines, [Chorus] with the hook, and a closing [Verse] or [Outro].
- Describe tempo feel, instruments and mood in Spanish.
- Write original lyrics rather than quoting a song; copyrighted lyrics are blocked.
- Keep each section short, since the Sume router rejects duration fields and the length is the model's choice.
- Use the optional image_url only when a still really sets the scene.
Checking the result on Sume
The Music Router docs say a music result can carry result.lyrics, a model-reported section map. Compare it with the lyrics you sent: if a section is missing or reordered, the track did not follow your structure and a re-run with a tighter prompt is the fix. Treat the map as the model's own report, not as a forced alignment. The lyrics section map post explains the difference.
Listen for pronunciation on names and borrowed words. Because there is no per-word control, the usual correction is to respell the word phonetically in the prompt.
| Step | Where it comes from | What to do |
|---|---|---|
| Prompt language | Google Lyria doc | Write prompt and lyrics in Spanish |
| Section tags | Google Lyria doc | Use [Verse], [Chorus] |
| Model id | Sume Music Router docs | lyria-3.5 or sume/music-auto |
| Structure check | Sume Music Router docs | Read result.lyrics |
Rights and labels
Lyria output carries a SynthID watermark, and Google blocks artist-voice requests. If you publish on YouTube, check whether the music is the main focus of the video, since the rules differ by use; see our YouTube label post. The same prompt discipline applies to any language: say what you want in the language you want it sung.
Sources
Related posts
More in Use cases
- A speaker withdrew voice consent: finding every narration that used it
Revocation is a clean-up job. Build a manifest of voice, model and job id for each narration, then regenerate with another voice and join audio on Sume.
- Splice an AI ending onto a Short and keep the original audio
Gemini Omni clips come with their own sound. To keep your Short's voice and music, detach both audio tracks and join them as audio parts in one Timeline render.
- Split a Short's audio at a cut point: timeline-audio ranges
YouTube Create has a split tool for audio. Sume's timeline-audio endpoint splits one file into up to 20 ranges, or joins up to 20 parts, for $0.01 a job.
- Sports replay Shorts: YouTube allows them if you explain the moves
YouTube lists sports replays where you explain what the competitor did as allowed. How to build a vertical breakdown Short from your own match footage.
Written by Sume