Lyria 3 Clip is fixed at 30 seconds: which Sume music id to use
Google's Lyria 3 Clip is a fixed 30-second MP3. Sume's Music Router offers lyria-3.5, lyria-3-pro and sume/music-auto, with length steered in the prompt.

Lyria 3 Clip is fixed at 30 seconds. Sume's Music Router does not list a clip id: its ids are sume/music-auto (the default, which resolves to lyria-3.5), lyria-3.5 and lyria-3-pro. For a 30-second bed, send sume/music-auto and ask for "a 30-second track" in the prompt. Each generation costs $0.125.
What Google says
The Gemini API music generation guide describes two models. Lyria 3 Clip (lyria-3-clip-preview) is fixed at 30 seconds, MP3. Lyria 3.5 (lyria-3.5) makes full songs of about a couple of minutes, controllable with the prompt, MP3 by default with WAV optional. Both are 44.1 kHz stereo and accept text plus up to 10 images. All audio carries a SynthID watermark, and results are non-deterministic.
What Sume accepts
Sume does not take a duration or duration_seconds field on music; the API rejects both. An optional single image_url can set the mood.
| Id | What it is |
|---|---|
| sume/music-auto | Default; resolves to lyria-3.5 |
| lyria-3.5 | Lyria 3.5, length steered in the prompt |
| lyria-3-pro | Lyria 3 Pro |
How to get a 30-second track
Put the length in the brief and name where the track changes. A 30-second bed has room for an intro, one lift, and an ending, so write it that way.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: bed-30s-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Bright acoustic pop, 100 BPM, G major. Fingerpicked guitar, soft claps, a glockenspiel lift at 0:15. A 30-second track. Instrumental, no vocals."
}'Check the length you got
Because length is steered by words, it is not exact. Read the duration from the artifact and, if the track runs long, trim it in Timeline with the soundtrack fade_out_seconds (up to 10) rather than regenerating. See exact length without a duration field for the trim route.
Because results are non-deterministic, two runs of the same prompt differ. If you need a choice, run a few and audition them.
Picking between the three ids
Leave the model as sume/music-auto unless you have a reason. It resolves to lyria-3.5 today, and job.request.routed_model tells you which engine ran. Choose lyria-3-pro when you specifically want that engine; the comparison post walks through the choice.
Whatever you pick, the price per generation is the fixed $0.125, so the choice is about sound, not cost.
Sources
Related posts
More in Models
- MAI-Transcribe-2-Streaming has no server VAD: who ends the turn?
In the preview, turn_detection only accepts null, so your client sends the commit event. Sume STT has no live socket; it cuts sentences after the job finishes.
- MAI-Voice-2.1 has 23 languages, 26 locales, 28 codes: which to quote
Microsoft says 23 languages and 26 locales; OpenRouter lists 28 codes and says 30+. Quote 23 languages, and use Python to turn the 28 codes into 23.
- MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS
Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.
- MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.
Written by Sume