AI music video from a song: Flow Music vs Lyria plus Timeline on Sume
Flow Music uses Gemini Omni to direct a music video by conversation. How to build a song with Lyria 3.5 and cut clips to it with Timeline 1.0 on Sume.

Flow Music lets you direct a music video by conversation, and on Sume you assemble one yourself: generate the song with the Music Router, generate 3-10 second clips with Gemini Omni Flash 1.1, and cut them over the song with Timeline 1.0. It is more manual than a chat, and it gives you an API and a file.
Google's I/O 2026 roundup says Flow Music uses Gemini Omni for conversational music video direction, with refinement of things like lyrics language, genre and instruments (read 2026-10-02). It does not say how long the video can be or whether it can be exported by API.
What does Sume do in each step?
Three Sume surfaces cover it. Each is an async job you poll.
| Step | Surface | Key limits from the docs |
|---|---|---|
| Song | Music Router POST /v1/music-router/generate | Prompt 1-5000 characters; length steered in the prompt; duration rejected |
| Clips | Video Router gemini-omni-flash-1.1 | 3-10 s each; 360p to 4K; 16:9 or 9:16; audio always on |
| Cut | Timeline 1.0 POST /v1/timeline-1.0/render | Audio 1-1800 s; 1-200 video slots; $0.10 per output minute |
How do I make the song?
The Music Router uses sume/music-auto, which is Lyria 3.5 today. Steer length and structure inside the prompt, for example "a 2-minute track" and section timestamps like [0:00-0:30] Intro. The docs say to put exclusions in the positive prompt, and negative_prompt is unsupported.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: song-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Dreamy synth-pop, 100 BPM. [0:00-0:20] Intro, soft pads. [0:20-1:00] Verse and chorus with female vocals in English. A 1-minute track."
}'How do I make clips that fit it?
Write one Omni prompt per section of the song: the intro clip, the chorus clip, and so on. Keep a recurring subject by reusing the same reference image as <IMAGE_REF_0> in each request. Clips are 3-10 seconds, so a 1-minute song needs roughly six to twelve of them, depending on how often you cut.
Omni clips come with generated audio that you will not want under your song. Timeline 1.0 takes a single audio spine and ordered video slots, so the clip audio is not used; the song is.
How do I cut them to the song?
Timeline takes audio.url (the song artifact, duration_seconds required) and a video[] list, each slot with source_url, start on the audio clock, and duration. All URLs must be Sume-hosted media.sume.com artifacts. Use plan first, which is unbilled and returns the cost estimate. Set transitions on slots after the first.
The result is one MP4 and comes back through GET /v1/jobs/:id/result as video_url. See Timeline 1.0 for fit modes, fps and refusals.
What does this not do?
Nothing in these docs syncs lip movement or on-screen action to lyrics or beats. You choose the cut points yourself, so listen to the song and place slot starts on the beats. And Sume has no conversational interface for changing the song after the fact: you re-run the Music Router with a revised prompt and rebuild the Timeline.
Keep the job ids from each step; see Jobs and results for recovering work after a restart.
Sources
Related posts
More in Use cases
- AI image with a working QR code: Muse Image uses code, others guess
Meta says Muse Image runs code to make accurate QR codes. Image models just draw one that may not scan. Generate the art, then add a real QR yourself.
- AI thumbnail variants in one call: n=4, pick one, then polish
Ask for four thumbnail options in one POST /v1/images call with n up to 4, pick the best by eye, then run a polish edit. Cost per step and when n is a waste.
- AI video with readable on-screen text: Kling 3.0 or Seedance 2.5
Kling pitches native text rendering; Dreamina pitches clean frames for text editing. What each page says and the Sume route that guarantees your exact words.
- AI-written script on a public-interest topic: Article 50 text label
Article 50(4) text labels cover published public-interest text with no human review or editorial control. What counts as review, per the Commission.
Written by Sume