Lip-sync a video you have, or animate a photo: which API?

Re-syncing a mouth in footage and making a still talk are different jobs. A guide using Sync.so docs and Sume lip-sync, avatar and face-swap routes.

5 min readSume
All posts

If you already have a video and want the speaker's mouth to match new audio, you want a video lip-sync API such as Sync.so. If you have only a photo or an avatar and a line of audio, you want a still-plus-audio route, which on Sume is the MiniMax H3 Max lip-sync endpoint or the avatar talking-video route. Choosing the wrong class is the most common reason a first test disappoints.

Sync.so's own docs list models named lipsync-1.9, lipsync-2, lipsync-2-pro, react-1 and sync-3, MP4 video with WAV or MP3 audio, and webhooks for async status (Sync Labs docs, read 2026-10-06). Sume's side comes from Models, Generate avatar video and Face swap (Beta), all read 2026-10-06.

Which input do you start from?

Start from what you can hold in your hand, not from the feature name.

Input-first routing. Sync.so details from https://sync.so/docs/introduction; Sume details from docs.sume.com, read 2026-10-06.
You haveJobSync.soSume
Footage of a speaker, new audioRe-sync the mouthYes, video plus audio inNot a lip-sync-on-video route; see face swap for a different face
A photo, a voice line of 5 to 14.8 sMake the still talkNot what its docs describePOST /v1/minimax/h3-max/lip-sync
An avatar and only a scriptPresenter clipNot what its docs describePOST /v1/avatar-1.0/talking-video, 4 to 60 s
Footage, but the person should be your avatarReplace the faceNot what its docs describeAvatar Face Swap (Beta), about 4 to 15 s with audio

What does Sume do with a video you already have?

Sume does not offer a route that re-syncs an arbitrary speaker's lips in your footage. The nearest routes change who is on screen. Face swap (Beta) applies a ready avatar's face to a public source video that has usable audio, roughly 4 to 15 seconds, and requires quality. H3 Max Recast, listed in the video catalog, replaces people in a 5 to 30 second source using 1 to 4 reference photos. Neither is a translation lip-sync.

Say so plainly in a vendor check: if the requirement is dubbing a real presenter into another language with matching lips, Sume is not the tool, and Sync.so or a similar service is. The model limits post explains why a general video model cannot do it either.

How should you test before you commit?

Run one sample of each class with the same line of audio. For the video route, use a clean 10-second clip with a front-facing speaker. For Sume's still route, use a front-facing still within the allowed aspect ratio, a Sume-hosted audio file of 5 to 14.8 seconds, and resolution: 768p. Watch teeth, consonants such as p and b, and what happens in the second after the audio ends.

Then check cost and failure handling, not just looks. Sume reserves at submit and refunds on failure; ask any vendor what happens to a rejected input. For volume, compare queueing: Sync.so's pricing page lists a Batch API on its Scale plan that can "launch 100s of videos with a single API call", while Sume takes one job per request with bulk runs described in the batch comparison.

Finally, write down the acceptance test before you look at results: for example, no visible lip drift after the third sentence, no dropped syllables at the end of the line, and a file you can play in your target player without re-encoding. A written bar keeps a pretty demo from outvoting a failing sample, and it makes the choice between a video re-sync and a still-plus-audio route a matter of evidence rather than taste.

How do you decide in practice?

Ask what you can reshoot. If the footage exists and only the words changed, a tool that re-times the mouth on video fits, and Sume's documented route is not that shape: it starts from a still or avatar. If you have no footage and want a presenter who can be re-scripted any time, a still plus audio is cheaper than a shoot.

Do not choose on price per second alone. Compare what the second contains: a rendered new clip against an edit of a clip you already own, and what each requires before it starts.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume