Lyria 3.5 result.lyrics is model-reported: check the song with STT

The lyrics field on a Sume music job is what the model reports, not a transcript. If the words matter, run the track through Sume STT at $0.01 a minute.

5 min readSume
All posts

On a finished Sume music job, result.lyrics carries the model-reported lyrics or section map when present. "Model-reported" means it is what the model says it wrote, not a transcript of what the audio contains. If the words matter, for captions or for a check on a sung language, transcribe the track with Sume STT at $0.01 per audio minute.

Where the field comes from

The music docs state that you read the audio artifact from result.artifacts[] where type is audio, and that result.lyrics carries model-reported lyrics or a section map when they are present. It is metadata that came back with the song. The router does not run recognition on the audio to produce it.

Google's Lyria guide says results are non-deterministic, so the same prompt can give a different track, and with it different words.

When it can differ from the audio

A mismatch is plausible wherever generative audio meets text. Sung words can be slurred or dropped, a section can be sung in a different order, and a model can report lines it planned but did not deliver. Treat the field as a draft until you have listened or transcribed.

  • Instrumental tracks should have no lyrics; if result.lyrics has words, listen.
  • A track in another language is worth a transcript before it goes into a campaign.
  • A track that will be captioned needs timings, which result.lyrics does not give.

Check it with STT

Sume STT 1.0 takes a public HTTPS audio_url and returns text, words[] and optional sentence segments[]. Send duration_seconds so the reservation matches the track; leaving it out reserves one minute. The maximum is 600 seconds.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-song-check-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/track.mp3",
    "language_code": "en",
    "duration_seconds": 120,
    "segmentation": { "mode": "sentence" }
  }'

Compare, then decide

Line up the recognised sentences against result.lyrics. If they agree, use the lyrics for display. If not, use the STT output, which comes from the audio. The same segments can feed an LRC file; see the LRC walkthrough.

Replace the demo URL with the audio artifact URL from your job. The $0.01 per minute rate is for the audio you send, so a 2-minute track costs about two cents to check.

What to do with a mismatch

If the transcript and result.lyrics differ, trust the transcript for captions, because it follows the audio. If the sung words are wrong for your use, run another take; results vary, and the next one may follow the brief better. Add the exact words you need to the prompt.

Spending two cents to check a track is cheap next to publishing the wrong line.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume