Lip-sync a video you have, or animate a photo: which API?
Re-syncing a mouth in footage and making a still talk are different jobs. A guide using Sync.so docs and Sume lip-sync, avatar and face-swap routes.

If you already have a video and want the speaker's mouth to match new audio, you want a video lip-sync API such as Sync.so. If you have only a photo or an avatar and a line of audio, you want a still-plus-audio route, which on Sume is the MiniMax H3 Max lip-sync endpoint or the avatar talking-video route. Choosing the wrong class is the most common reason a first test disappoints.
Sync.so's own docs list models named lipsync-1.9, lipsync-2, lipsync-2-pro, react-1 and sync-3, MP4 video with WAV or MP3 audio, and webhooks for async status (Sync Labs docs, read 2026-10-06). Sume's side comes from Models, Generate avatar video and Face swap (Beta), all read 2026-10-06.
Which input do you start from?
Start from what you can hold in your hand, not from the feature name.
| You have | Job | Sync.so | Sume |
|---|---|---|---|
| Footage of a speaker, new audio | Re-sync the mouth | Yes, video plus audio in | Not a lip-sync-on-video route; see face swap for a different face |
| A photo, a voice line of 5 to 14.8 s | Make the still talk | Not what its docs describe | POST /v1/minimax/h3-max/lip-sync |
| An avatar and only a script | Presenter clip | Not what its docs describe | POST /v1/avatar-1.0/talking-video, 4 to 60 s |
| Footage, but the person should be your avatar | Replace the face | Not what its docs describe | Avatar Face Swap (Beta), about 4 to 15 s with audio |
What does Sume do with a video you already have?
Sume does not offer a route that re-syncs an arbitrary speaker's lips in your footage. The nearest routes change who is on screen. Face swap (Beta) applies a ready avatar's face to a public source video that has usable audio, roughly 4 to 15 seconds, and requires quality. H3 Max Recast, listed in the video catalog, replaces people in a 5 to 30 second source using 1 to 4 reference photos. Neither is a translation lip-sync.
Say so plainly in a vendor check: if the requirement is dubbing a real presenter into another language with matching lips, Sume is not the tool, and Sync.so or a similar service is. The model limits post explains why a general video model cannot do it either.
How should you test before you commit?
Run one sample of each class with the same line of audio. For the video route, use a clean 10-second clip with a front-facing speaker. For Sume's still route, use a front-facing still within the allowed aspect ratio, a Sume-hosted audio file of 5 to 14.8 seconds, and resolution: 768p. Watch teeth, consonants such as p and b, and what happens in the second after the audio ends.
Then check cost and failure handling, not just looks. Sume reserves at submit and refunds on failure; ask any vendor what happens to a rejected input. For volume, compare queueing: Sync.so's pricing page lists a Batch API on its Scale plan that can "launch 100s of videos with a single API call", while Sume takes one job per request with bulk runs described in the batch comparison.
Finally, write down the acceptance test before you look at results: for example, no visible lip drift after the third sentence, no dropped syllables at the end of the line, and a file you can play in your target player without re-encoding. A written bar keeps a pretty demo from outvoting a failing sample, and it makes the choice between a video re-sync and a still-plus-audio route a matter of evidence rather than taste.
How do you decide in practice?
Ask what you can reshoot. If the footage exists and only the words changed, a tool that re-times the mouth on video fits, and Sume's documented route is not that shape: it starts from a still or avatar. If you have no footage and want a presenter who can be re-scripted any time, a still plus audio is cheaper than a shoot.
Do not choose on price per second alone. Compare what the second contains: a rendered new clip against an edit of a clip you already own, and what each requires before it starts.
Sources
Related posts
More in Comparisons
- LTX-2.5 or a lip-sync API for a talking head: what Sume offers
Sume does not list LTX-2.5. For a talking head that must say your exact words, the route that ships is TTS plus H3 Max lip sync, or Avatar 1.0 talking video.
- MAI-Transcribe-2 at $0.54 an hour vs Sume STT at 1 cent a minute
MAI-Transcribe-2-Streaming is $0.54 an hour as an intro price through 2026. Sume STT is $0.01 a minute, $0.60 an hour. The bill at 10, 100 and 1,000 hours.
- MAI-Voice-2.1 or Sume TTS? Three questions that decide it
Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.
- MiniMax H3 open weights vs a hosted lip-sync API: what you take on
MiniMax released H3 open weights on 2026-08-03. Self-hosting is not the same as a hosted still-plus-audio lip-sync route. What each choice makes you own.
Written by Sume