Lyria 3.5 has no edit pass: iterate a music bed for $0.125 a take
Google says Lyria 3.5 is single-turn. Iterate a bed on Sume by changing one prompt axis per take; five takes cost $0.625 and a retry needs a key.

Lyria 3.5 cannot be edited after it returns, so iteration means generating again. Google's music generation page, read on 2026-10-05, describes the model as single-turn, with no iterative editing of a finished track. On Sume each accepted generation is a fixed $0.125 through the Music Router, so five takes cost $0.625 and a cap is easy to set.
A disciplined loop
- Take 1: your full brief, written with the seven-axis template.
- Takes 2 to 5: change exactly one axis per take and keep everything else identical, so you know what caused the difference.
- Stop at five. If none works, the brief is wrong, not the dice.
- Use a distinct Idempotency-Key per take, and reuse a key only to retry the same take after a network error.
What not to rely on
The router rejects duration and duration_seconds, and a non-empty negative_prompt returns 400 negative_prompt_unsupported, so exclusions go in the prompt text: "no vocals, no drums". The default sume/music-auto resolves to Lyria 3.5 today and may move without notice; pin lyria-3.5 while you iterate so takes are comparable. job.request.routed_model in the job record shows which engine ran, as explained in Jobs and results.
For reference, Google's pricing page, read on 2026-10-05, lists Lyria 3.5 at $0.08 per song with no free tier, while Sume's $0.125 is the price on a Sume account.
Sources
Related posts
More in Models
- Every Lyria 3.5 track carries SynthID: what brands should know
Google says all Lyria 3.5 output carries a SynthID audio watermark and blocks artist voices and copyrighted lyrics. Here is what that means for brand music.
- MAI-Transcribe-2-Streaming has no server VAD: who ends the turn?
In the preview, turn_detection only accepts null, so your client sends the commit event. Sume STT has no live socket; it cuts sentences after the job finishes.
- MAI-Voice-2.1 has 23 languages, 26 locales, 28 codes: which to quote
Microsoft says 23 languages and 26 locales; OpenRouter lists 28 codes and says 30+. Quote 23 languages, and use Python to turn the 28 codes into 23.
- MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS
Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.
Written by Sume