AI music video with a singer close-up: still plus audio, 5 to 14.8 s

The lip-sync model in Sume's video catalog takes a still and audio of 5 to 14.8 seconds. Split a song to fit, and join the shots on a Timeline.

4 min readSume
All posts

To make a singer close-up for an AI music video, the lip-sync entry in Sume's catalog takes one still image and one audio clip between 5 and 14.8 seconds long. A full song is far longer, so you cut the vocal into pieces that fit, make one shot per piece and rejoin them in order.

The constraint that shapes the plan

The limit comes from the model's own entry in the Sume docs, which lists the still and audio as the inputs and 5 to 14.8 seconds as the allowed audio length. Anything outside that range needs to be trimmed or padded first.

Lip-sync inputs and the pieces a song needs, from the Sume docs, read 2026-10-06.
ItemRuleConsequence
ImageOne still of the singerUse a clean, front-facing frame
Audio5 to 14.8 sSplit the vocal at phrase ends
Per shotOne call eachA 3-minute song is at least 13 shots
JoinTimeline 1.0 audio spinePlace each shot against the full track

Splitting the vocal

Cut at the end of a phrase, not mid-word, and aim for pieces comfortably inside the range, such as 8 to 12 seconds. Sume's Timeline audio tools can split audio, so the cut can happen on the platform, and the docs page for audio tools is where to check what each operation costs.

Keep the same still for every shot, so the singer looks the same throughout. Vary the picture by mixing in B-roll from the ordinary video models between the close-ups.

Assembly

Put the full track on the Timeline as the audio spine, then drop each shot where its audio piece begins. Because every shot was made from the real audio, the lips should match where they land. Review each join, and regenerate only the shot that looks off.

Check the licence on the song before you publish. Sume does not clear music rights for you.

Pick the still with care

The still sets the singer's look for the whole video, so spend your time there. Choose an image with the mouth clearly visible, even light and no hand across the face. A singer with the mouth half hidden gives the model little to work with, and every one of your shots inherits the problem.

Sources

Related posts

More in Models

All Models posts

Written by Sume