Lyria 3.5 tempo and duration control: Flow Music versus the Sume API
Google says Lyria 3.5 gives easier tempo and duration control in Flow Music. On Sume the same model has no duration field, so you steer length in the prompt.

Google's Lyria 3.5 announcement lists creative control as one of four advances: you can more easily control the tempo and duration of your outputs in Flow Music. Sume's Music 1.0 runs on Lyria 3.5 too, but it has no duration or tempo parameter. Tempo and length live in the prompt, and a request that sends duration or duration_seconds is rejected.
Where the control lives
The Flow Music announcement describes an app feature, and the blog post does not say how the control is exposed or whether it is in an API. So do not assume a field you saw in an app exists on a REST route. In Sume the request has prompt (1 to 5000 characters), an optional image_url, metadata, and the usual mode and webhook_url fields.
What to write in the prompt
Write the tempo as a number, such as "72 BPM", and the length as a plain phrase, such as "a 2-minute track", or use section markers like [0:00-0:30] Intro: .... The Sume docs describe these as creative directions, not guaranteed output values, so check the length of the file you get back.
curl -X POST https://api.sume.com/v1/music-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: score-72bpm-001" \
-d '{"prompt": "A 30-second neo-soul nocturne, 72 BPM, D minor. Rhodes and soft sub bass. Instrumental, no vocals."}'How the two surfaces compare, read 2026-10-06:
| Question | Flow Music (Google blog) | Sume Music 1.0 (docs) |
|---|---|---|
| Model | Lyria 3.5 | Lyria 3.5 |
| Tempo control | Easier control, per Google | BPM written in the prompt |
| Duration control | Easier control, per Google | No field; named in the prompt |
| Negative prompt | Not stated | Non-empty value returns 400 |
| Price | Not stated | $0.125 per generation |
Checking what you got
Because the prompt only steers the result, measure it. Read the duration of the returned file, and tap out the tempo against a metronome or ask your editor to show the BPM. If the length is wrong, do not rerun at random. Change one phrase of the prompt, such as the section markers, and compare.
Provider lyrics on the result describe tempo and structure, but the docs call them model-reported metadata, not a measurement of the audio. Treat the file as the source of truth, and keep the take that fits the cut.
If the length must be exact, generate a bit long and cut it. A timeline render takes the music as a soundtrack with a fade_out_seconds of up to 10, so the output ends where the picture ends. Confirm the live rate in GET /v1/catalog.
Sources
Related posts
More in Models
- Lyria 3.5 better vocals and pronunciation: a listening checklist
Google says Lyria 3.5 improves vocals, lyrics and pronunciation. Check a Sume Music 1.0 take yourself: listen, transcribe it with STT, and compare the words.
- MAI-Voice-2.1 vs Flash: what $7 per million saves on a 60-second short
Flash lists at $15 per million characters and MAI-Voice-2.1 at $22. On a 750-character short that is a fraction of a cent. Math for 1, 30 and 3,000 shorts.
- MCP server for video analysis: what can an agent actually read?
Sume's remote MCP lets an agent probe a clip, sample stills, pull exact frames and transcribe audio. Semantic scene tools are dev-only. Limits and costs inside.
- MiniMax H3 first request: a 5-second 768p clip for 38 cents on Sume
A 5-second 768p MiniMax H3 clip costs $0.38 on Sume, with stereo sound included. The request, the 15-second limit and the 2K and 4K upscale prices.
Written by Sume