Lyria 3.5 better vocals and pronunciation: a listening checklist
Google says Lyria 3.5 improves vocals, lyrics and pronunciation. Check a Sume Music 1.0 take yourself: listen, transcribe it with STT, and compare the words.

Google's Lyria 3.5 post says the model generates higher quality lyrics with improved prompt adherence and structural awareness, and more realistic, emotionally nuanced vocals with improved pronunciation. Those are the vendor's claims for Flow Music. Sume's Music 1.0 runs on Lyria 3.5, so the same improvements may apply, but the way to know is to test a take. You can do that with one generation and one transcription.
Generate a take
Write the lyrics you want sung into the prompt, with section markers like [0:00-0:30] Intro: ... if you want structure. Sume's docs say the provider lyrics field is model-reported metadata and not an audio measurement, so it does not prove what was sung. Treat it as a hint and check the audio.
Transcribe it back
Send the finished track's URL to POST /v1/stt-1.0/transcribe with the language you expect. Compare the returned text with the lyrics you wrote. Words that differ are either a recognizer miss or a sung word that was changed. Listen to those places to tell which.
- Does every line you wrote appear, in order?
- Are names and numbers pronounced as written?
- Is the vocal style the one you asked for, or the default?
- Does the track match your named tempo and length?
- Is the audio free of words you did not write?
Keep your own notes per take. Each generation is billed once, so a rerun costs the same as the first take. What the vendor claims, and what you check, read 2026-10-06:
| Claim (Google blog) | Check on Sume |
|---|---|
| Higher quality lyrics | Compare STT text with your lyrics |
| More realistic vocals | Listen for the style you named |
| Improved pronunciation | Spot-check names and numbers |
| Easier tempo and duration control | Measure the file length |
Keeping a record of takes
Save the prompt, the job id and the STT text next to each take. If a client asks why one version was chosen, you can show the lyrics you asked for, what was sung and what you picked.
A published vocal track may need a disclosure that it is synthetic, depending on the platform. Check the platform's own rules before you upload, and do not publish a track that sings a name or a claim you did not write.
Music 1.0 costs $0.125 per generation at the public rate. Confirm the live price in GET /v1/catalog, and do not publish a vocal track before you have listened to the whole file.
Sources
Related posts
More in Models
- MAI-Voice-2.1 vs Flash: what $7 per million saves on a 60-second short
Flash lists at $15 per million characters and MAI-Voice-2.1 at $22. On a 750-character short that is a fraction of a cent. Math for 1, 30 and 3,000 shorts.
- MCP server for video analysis: what can an agent actually read?
Sume's remote MCP lets an agent probe a clip, sample stills, pull exact frames and transcribe audio. Semantic scene tools are dev-only. Limits and costs inside.
- MiniMax H3 first request: a 5-second 768p clip for 38 cents on Sume
A 5-second 768p MiniMax H3 clip costs $0.38 on Sume, with stereo sound included. The request, the 15-second limit and the 2K and 4K upscale prices.
- MiniMax H3 open weights: 768p locally, 2K only hosted
MiniMax's open H3 Base checkpoints top out at 768p. The 2K tier needs a proprietary module. What that means if you want to self-host or call H3 through Sume.
Written by Sume