Which language does Lyria 3.5 sing in? Spanish prompt on Sume Music

Google says Lyria generates music in the language of your prompt. On Sume, write the brief in the language you want sung; no language field exists on Music.

4 min readSume
All posts

Lyria generates music in the language of your prompt, per Google's documentation, so a Spanish brief is how you ask for Spanish vocals. Sume's Music request has no language field, which means the prompt itself is the only language control you have.

The Lyria statement comes from the Gemini API music page (read 2026-10-10). The Sume request shape comes from the Music 1.0 page: prompt, optional image_url, optional metadata, and the delivery fields.

What the same page rules out

The same Google page says safety filters check all prompts and that requests for specific artist voices or copyrighted lyrics are blocked. So 'sing the chorus of a famous song in Spanish' is the wrong brief. Describe a style, a mood and an original line instead.

Sume's docs say that after a policy rejection you should change the flagged content but keep the musical brief, and try again only within the authorized budget.

A brief that asks for Spanish vocals

Write the whole brief in Spanish, or write the English brief and quote the lyric line in Spanish. Google documents only that the language follows the prompt; it does not promise pronunciation quality for every language, so check the first take by ear before you commit to a batch.

Where language is set (Lyria page read 2026-10-10; Sume per docs.sume.com)
SurfaceWhere you set the sung language
Lyria through the Gemini APIIn the prompt; output follows the language of the prompt
Sume Music 1.0 and Music RouterIn the prompt; there is no language field
Sume TTS (for comparison)A language field; omitted means English

Image conditioning and language

Sume's Music request also accepts one optional public HTTPS image_url, and Google's page says Lyria accepts up to 10 images with text. An image changes the mood, not the sung language, so keep the language instruction in the text prompt. If you attach a scene still, the Music page suggests passing the accepted scene still as image_url when you want continuity across a project.

Remember the prompt cap of 5,000 characters. A full brief with seven axes and a quoted lyric in another language is well under it, but a pasted script is not. Keep lyrics to the lines you intend to be sung and describe everything else in a few sentences.

Do not mix this up with TTS

Text-to-speech on Sume is different: its request has a language field and the OpenAPI contract says to set it for every non-English transcript, because an omitted value defaults to English at the provider. Music has no such field. If you want spoken Spanish narration over an instrumental bed, make two files: a TTS job with language set to es, and a Music job with 'Instrumental, no vocals' in the prompt. The Music page recommends ending an instrumental brief with that clause.

See setting the TTS language field for the speech side.

Practical steps

A short routine keeps costs predictable at a fixed $0.125 per generation:

Finally, decide where the spoken or sung line sits in the final video. If the sung line must land on a specific scene cut, generate the track first, measure the real duration of the file, and then place the cuts on a Timeline render against that measured length rather than the length you asked for in the prompt.

  • Write the brief in the target language, with tempo as a number and an instrument list.
  • Keep any lyric line original and short; avoid quoting existing songs.
  • Add 'no spoken word' only if you will lay narration over it.
  • Listen for the sung language before committing; budget two or three takes, $0.25 to $0.375.
  • Store the file and the exact prompt together.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume