Lyria 3.5 result.lyrics is model-reported: check the song with STT
The lyrics field on a Sume music job is what the model reports, not a transcript. If the words matter, run the track through Sume STT at $0.01 a minute.

On a finished Sume music job, result.lyrics carries the model-reported lyrics or section map when present. "Model-reported" means it is what the model says it wrote, not a transcript of what the audio contains. If the words matter, for captions or for a check on a sung language, transcribe the track with Sume STT at $0.01 per audio minute.
Where the field comes from
The music docs state that you read the audio artifact from result.artifacts[] where type is audio, and that result.lyrics carries model-reported lyrics or a section map when they are present. It is metadata that came back with the song. The router does not run recognition on the audio to produce it.
Google's Lyria guide says results are non-deterministic, so the same prompt can give a different track, and with it different words.
When it can differ from the audio
A mismatch is plausible wherever generative audio meets text. Sung words can be slurred or dropped, a section can be sung in a different order, and a model can report lines it planned but did not deliver. Treat the field as a draft until you have listened or transcribed.
- Instrumental tracks should have no lyrics; if
result.lyricshas words, listen. - A track in another language is worth a transcript before it goes into a campaign.
- A track that will be captioned needs timings, which
result.lyricsdoes not give.
Check it with STT
Sume STT 1.0 takes a public HTTPS audio_url and returns text, words[] and optional sentence segments[]. Send duration_seconds so the reservation matches the track; leaving it out reserves one minute. The maximum is 600 seconds.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-song-check-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/track.mp3",
"language_code": "en",
"duration_seconds": 120,
"segmentation": { "mode": "sentence" }
}'Compare, then decide
Line up the recognised sentences against result.lyrics. If they agree, use the lyrics for display. If not, use the STT output, which comes from the audio. The same segments can feed an LRC file; see the LRC walkthrough.
Replace the demo URL with the audio artifact URL from your job. The $0.01 per minute rate is for the audio you send, so a 2-minute track costs about two cents to check.
What to do with a mismatch
If the transcript and result.lyrics differ, trust the transcript for captions, because it follows the audio. If the sung words are wrong for your use, run another take; results vary, and the next one may follow the brief better. Add the exact words you need to the prompt.
Spending two cents to check a track is cheap next to publishing the wrong line.
Sources
Related posts
More in Media tools
- Turn a Lyria 3.5 MP3 into a WAV with one Timeline audio job
Sume's music jobs usually return MP3. A single-part Timeline audio concat can output sample-exact WAV for $0.01, which matters before you cut or re-join it.
- Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file
A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.
- Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.
- Microsoft's Content Provenance Detection: what you can check on a file
Foundry has a detection website and API for provenance. What the page says it checks, its limits, and how to use it on an AI clip or image from any generator.
Written by Sume