AI music video from a song: Flow Music vs Lyria plus Timeline on Sume

Flow Music uses Gemini Omni to direct a music video by conversation. How to build a song with Lyria 3.5 and cut clips to it with Timeline 1.0 on Sume.

5 min readSume
All posts

Flow Music lets you direct a music video by conversation, and on Sume you assemble one yourself: generate the song with the Music Router, generate 3-10 second clips with Gemini Omni Flash 1.1, and cut them over the song with Timeline 1.0. It is more manual than a chat, and it gives you an API and a file.

Google's I/O 2026 roundup says Flow Music uses Gemini Omni for conversational music video direction, with refinement of things like lyrics language, genre and instruments (read 2026-10-02). It does not say how long the video can be or whether it can be exported by API.

What does Sume do in each step?

Three Sume surfaces cover it. Each is an async job you poll.

Music video pipeline on Sume, from Sume docs, read 2026-10-02
StepSurfaceKey limits from the docs
SongMusic Router POST /v1/music-router/generatePrompt 1-5000 characters; length steered in the prompt; duration rejected
ClipsVideo Router gemini-omni-flash-1.13-10 s each; 360p to 4K; 16:9 or 9:16; audio always on
CutTimeline 1.0 POST /v1/timeline-1.0/renderAudio 1-1800 s; 1-200 video slots; $0.10 per output minute

How do I make the song?

The Music Router uses sume/music-auto, which is Lyria 3.5 today. Steer length and structure inside the prompt, for example "a 2-minute track" and section timestamps like [0:00-0:30] Intro. The docs say to put exclusions in the positive prompt, and negative_prompt is unsupported.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: song-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Dreamy synth-pop, 100 BPM. [0:00-0:20] Intro, soft pads. [0:20-1:00] Verse and chorus with female vocals in English. A 1-minute track."
  }'

How do I make clips that fit it?

Write one Omni prompt per section of the song: the intro clip, the chorus clip, and so on. Keep a recurring subject by reusing the same reference image as <IMAGE_REF_0> in each request. Clips are 3-10 seconds, so a 1-minute song needs roughly six to twelve of them, depending on how often you cut.

Omni clips come with generated audio that you will not want under your song. Timeline 1.0 takes a single audio spine and ordered video slots, so the clip audio is not used; the song is.

How do I cut them to the song?

Timeline takes audio.url (the song artifact, duration_seconds required) and a video[] list, each slot with source_url, start on the audio clock, and duration. All URLs must be Sume-hosted media.sume.com artifacts. Use plan first, which is unbilled and returns the cost estimate. Set transitions on slots after the first.

The result is one MP4 and comes back through GET /v1/jobs/:id/result as video_url. See Timeline 1.0 for fit modes, fps and refusals.

What does this not do?

Nothing in these docs syncs lip movement or on-screen action to lyrics or beats. You choose the cut points yourself, so listen to the song and place slot starts on the beats. And Sume has no conversational interface for changing the song after the fact: you re-run the Music Router with a revised prompt and rebuild the Timeline.

Keep the job ids from each step; see Jobs and results for recovering work after a restart.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume