Every Lyria 3.5 track carries SynthID: what brands should know
Google says all Lyria 3.5 output carries a SynthID audio watermark and blocks artist voices and copyrighted lyrics. Here is what that means for brand music.

Google's Gemini API guide says a SynthID audio watermark is applied to all Lyria output. For a brand, that means any track you generate can later be identified as AI-generated by tools that read SynthID. It is not a licence, and it does not make a track safe to use. The guide also says prompts that request specific artist voices or copyrighted lyrics are blocked.
Here is what the guide states, what it does not, and how this applies if you reach Lyria 3.5 through Sume.
What Google's page says
All rows below are from the Gemini API music generation guide, read 2026-10-05.
| Topic | What the guide states |
|---|---|
| Watermark | SynthID audio watermark on all output |
| Artist voices | Prompts requesting specific artist voices are blocked |
| Lyrics | Prompts requesting copyrighted lyrics are blocked |
| Output format | 44.1 kHz stereo, MP3 or WAV |
| Models | lyria-3-clip-preview for 30-second clips; lyria-3.5 for full songs of a couple of minutes |
| Repeatability | Single-turn only; identical calls give varying outputs |
What a watermark does and does not do
A watermark is a provenance signal. It travels with the audio so a platform or a partner can check whether Google's system made it. For brands that publish on channels that ask for AI disclosure, that is useful: you can answer the question truthfully because the signal is already in the file.
It does not tell anyone who owns the track, whether your use is licensed, or whether the melody resembles existing music. The guide does not make those claims, and neither should you. If a contract needs warranties about the music, ask the supplier, not the watermark.
- Plan on disclosure. Treat generated music as detectable as AI-made.
- Do not try to prompt around the artist and lyrics blocks. The guide says those requests are refused, and a block is a sign to rewrite the brief.
- Keep the exact prompt and job record with the delivered file, because outputs vary and you cannot regenerate the same track.
If you use Lyria through Sume
Sume's Music 1.0 runs on Lyria 3.5, and the Music Router routes sume/music-auto to it today. Sume's docs do not mention SynthID, so I cannot state from them that the watermark is present in Sume-delivered files. Google's statement covers its own Gemini API output; confirm with a test file and your own detector before you promise a client either way.
What the Sume docs do say: results are hosted audio artifacts on media.sume.com, job.request.routed_model names the engine that ran, and an explicit lyria-3.5 id pins it. Those fields are your audit trail. The docs also state that the policy rejection path is to change the flagged content while keeping the musical brief.
A short brand checklist
Write the brief around mood, tempo, key and two to four instruments, never around a named artist. Save the prompt, job id and routed model next to the file. Ask your legal team to read Google's current terms. If an audience-facing platform needs an AI label, add it. None of these steps depend on the watermark, which is one more piece of evidence rather than the whole answer.
Sources
Related posts
More in Models
- MAI-Transcribe-2-Streaming has no server VAD: who ends the turn?
In the preview, turn_detection only accepts null, so your client sends the commit event. Sume STT has no live socket; it cuts sentences after the job finishes.
- MAI-Voice-2.1 has 23 languages, 26 locales, 28 codes: which to quote
Microsoft says 23 languages and 26 locales; OpenRouter lists 28 codes and says 30+. Quote 23 languages, and use Python to turn the 28 codes into 23.
- MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS
Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.
- MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.
Written by Sume