Ten holiday ad tracks: Lyria 3.5 at Google's list vs Sume Music
Google lists Lyria 3.5 at $0.08 a song and Sume Music bills $0.125 a generation, so ten holiday ad tracks cost $0.80 or $1.25. Here is what the gap covers.

Ten holiday ad tracks cost $0.80 at Google's Lyria 3.5 list price and $1.25 on Sume Music, a difference of $0.45. For a campaign that is a rounding error next to the media spend, but it is worth knowing what it buys.
Google's Gemini API pricing page, read on 2026-10-04, lists Lyria 3.5 at $0.08 per song. Sume's Music Router uses Lyria 3.5 as its current engine and bills $0.125 per audio generation.
The ten-track budget
| Tracks | Google at $0.08 | Sume at $0.125 |
|---|---|---|
| 1 | $0.08 | $0.125 |
| 10 | $0.80 | $1.25 |
| 25 | $2.00 | $3.125 |
| 100 | $8.00 | $12.50 |
What the Sume price includes
On Sume the track is a job like any other: it lands as a hosted audio asset you can pass straight to a Timeline render, a caption job or an avatar video. The prompt is 1 to 5,000 characters, and the docs say a duration field is rejected, so you set length in the prompt itself, for example by asking for a 2-minute track or by writing timed sections.
If you generate in a script against Google's API directly, you pay less per song and handle storage and hand-off yourself. If you only need a few tracks for one campaign, the $0.45 difference does not decide anything, and convenience can.
A practical order of work
- Write ten short briefs, one per ad cut, with the mood and the tempo you want.
- Generate all ten, listen, and keep the best three.
- Reserve a retry or two per track in your budget, since taste is not deterministic.
Sources
Related posts
More in Comparisons
- Image reference limits: 10 on ElevenLabs, 16 on Sume, 5 Ideogram 4.5
ElevenLabs lists 10 references for GPT Image 2.5. Sume's docs list 16 for the same models and 5 total for Ideogram 4.5. A table, a count guard and an edit rule.
- gpt-live-transcribe has no word timestamps: what it means for captions
OpenAI's realtime guide says gpt-live-transcribe cannot return word-level timestamps. If you need karaoke captions, that decides which path to use.
- Grok Imagine Video 1.5: Runway's listing vs Sume's catalog row
Sume's grok-imagine-video-1.5 row is image-to-video only, 480p or 720p, 4 to 15 s, no audio flag. Runway's changelog lists text-to-video and audio.
- Grok STT smart_turn end-of-turn vs Sume STT sentence segmentation
smart_turn predicts when a speaker is done in live streams. Sume STT is batch and cuts sentences from word timings. Which one do you need?
Written by Sume