Dub with lip sync: Meta Reels option vs Sume Avatar 1.0 (English-only)
Meta offers optional lip sync on translated Reels. Sume Avatar 1.0 is English-only, so a non-English talking shot uses TTS plus a lip-sync endpoint.
On Meta's side, lip sync is an option a creator turns on for a translated Reel: Meta's newsroom post on Reels translation says it syncs the translated audio to the creator's mouth movements (read 2026-10-02). On Sume's side, Avatar 1.0 talking video is English-only, so a talking shot in Spanish, Korean or any other language does not go through Avatar 1.0. Make the voice with TTS in that language, then animate a still with a lip-sync endpoint.
The two are different jobs. Meta's feature dubs a video that already exists. Sume's lip-sync endpoints build a talking clip from a still and audio you supply.
What does each option take as input?
Meta's tool takes a Reel you published. Sume's Avatar 1.0 takes an avatar and a script, and speaks English. The still-plus-audio endpoints take Sume-hosted audio instead of a script, so the language comes from the audio you made.
The Models overview lists the two lip-sync surfaces; H3 Max takes audio of 5 to 14.8 seconds.
| Option | Input | Language |
|---|---|---|
| Meta AI Reels translation | A Reel you posted; lip sync is optional | The languages Meta lists on its page |
| Sume Avatar 1.0 talking video | An avatar plus a script | English only |
| VEED Fabric 1.0 | A still plus Sume-hosted audio | Set by the audio |
| MiniMax H3 Max Lip Sync | The same still plus audio body, audio 5 to 14.8 s | Set by the audio |
How do I make a non-English talking shot?
Generate the voice first with POST /v1/tts-1.0/generate and set language to the language of the script; the OpenAPI says to set it for every non-English transcript because an omitted value defaults to English at the provider. Then send the resulting audio URL and your still to the Fabric endpoint.
The Sume docs do not claim an avatar's own voice speaks other languages, so listen to the output before you commit a batch.
Which should I choose?
If the video is already on Instagram or Facebook and the language is on Meta's list, use the platform option. If you need a talking shot in a language Avatar 1.0 does not speak, build it from TTS and a lip-sync endpoint. For the longer comparison of dubbing and lip sync, see lip sync vs dubbing.
Sources
Related posts
More in Sume Avatar 1.0
- Regenerate avatar preview stills, or start a new preview?
Regenerate refreshes first-frame stills from the stored preview request. Changing script, avatar, scene or aspect ratio needs a new preview. The full rule.
- One avatar handle, three platform cuts: keep the character consistent
Keep one presenter across LinkedIn, Snapchat and Pinterest by reusing a single avatar_handle and rendering each cut at the right ratio and length with Sume.
- Snapchat Spotlight needs 6 seconds: avatar clips that start at 4
Snap's Public Profile API takes Spotlight videos of 6 to 60 seconds, but Avatar 1.0 plans from 4. Plan scripts at 6 seconds or more and probe before posting.
- Sume Avatar API: canonical routes vs the legacy model-run aliases
Which Avatar 1.0 endpoint should a new integration call? The canonical /v1/avatar-1.0 routes, with the legacy aliases kept for compatibility. All paths listed.
Written by Sume