AI music video with a singer close-up: still plus audio, 5 to 14.8 s
The lip-sync model in Sume's video catalog takes a still and audio of 5 to 14.8 seconds. Split a song to fit, and join the shots on a Timeline.

To make a singer close-up for an AI music video, the lip-sync entry in Sume's catalog takes one still image and one audio clip between 5 and 14.8 seconds long. A full song is far longer, so you cut the vocal into pieces that fit, make one shot per piece and rejoin them in order.
The constraint that shapes the plan
The limit comes from the model's own entry in the Sume docs, which lists the still and audio as the inputs and 5 to 14.8 seconds as the allowed audio length. Anything outside that range needs to be trimmed or padded first.
| Item | Rule | Consequence |
|---|---|---|
| Image | One still of the singer | Use a clean, front-facing frame |
| Audio | 5 to 14.8 s | Split the vocal at phrase ends |
| Per shot | One call each | A 3-minute song is at least 13 shots |
| Join | Timeline 1.0 audio spine | Place each shot against the full track |
Splitting the vocal
Cut at the end of a phrase, not mid-word, and aim for pieces comfortably inside the range, such as 8 to 12 seconds. Sume's Timeline audio tools can split audio, so the cut can happen on the platform, and the docs page for audio tools is where to check what each operation costs.
Keep the same still for every shot, so the singer looks the same throughout. Vary the picture by mixing in B-roll from the ordinary video models between the close-ups.
Assembly
Put the full track on the Timeline as the audio spine, then drop each shot where its audio piece begins. Because every shot was made from the real audio, the lips should match where they land. Review each join, and regenerate only the shot that looks off.
Check the licence on the song before you publish. Sume does not clear music rights for you.
Pick the still with care
The still sets the singer's look for the whole video, so spend your time there. Choose an image with the mouth clearly visible, even light and no hand across the face. A singer with the mouth half hidden gives the model little to work with, and every one of your shots inherits the problem.
Sources
Related posts
More in Models
- AI video ad text in 12 languages: Wan 3.0 or burned-in captions?
Alibaba says Wan 3.0 renders text in 12 languages. Sume lists wan-3.0 at 2 to 30 s. When to let the model draw words and when to add captions yourself.
- AI video API news, October 6 2026: what changed for callers
Kling 4.0 heads to full launch, Seedance 2.5 API pages disagree, Omni 1.1 Flash is Google's default. What to change in a video integration today.
- AI video generator for long videos: 30 seconds per pass, then assemble
ByteDance says Seedance 2.5 makes 30 seconds a pass and extends to minutes. On Sume, a video job tops out at 30 s; Timeline 1.0 joins clips to 1800 s.
- Bumper sticker art: a 3:1 strip at 1536x512 on GPT Image 2.5
Make a wide bumper sticker on Sume with GPT Image 2.5 at a custom 1536x512 size, why 3:1 is its widest shape, and what to do for longer stickers.
Written by Sume