Dub with lip sync: Meta Reels option vs Sume Avatar 1.0 (English-only)

Meta offers optional lip sync on translated Reels. Sume Avatar 1.0 is English-only, so a non-English talking shot uses TTS plus a lip-sync endpoint.

5 min readSume
All posts

On Meta's side, lip sync is an option a creator turns on for a translated Reel: Meta's newsroom post on Reels translation says it syncs the translated audio to the creator's mouth movements (read 2026-10-02). On Sume's side, Avatar 1.0 talking video is English-only, so a talking shot in Spanish, Korean or any other language does not go through Avatar 1.0. Make the voice with TTS in that language, then animate a still with a lip-sync endpoint.

The two are different jobs. Meta's feature dubs a video that already exists. Sume's lip-sync endpoints build a talking clip from a still and audio you supply.

What does each option take as input?

Meta's tool takes a Reel you published. Sume's Avatar 1.0 takes an avatar and a script, and speaks English. The still-plus-audio endpoints take Sume-hosted audio instead of a script, so the language comes from the audio you made.

The Models overview lists the two lip-sync surfaces; H3 Max takes audio of 5 to 14.8 seconds.

Inputs compared from Meta's newsroom page and the Sume Models overview, read 2026-10-02.
OptionInputLanguage
Meta AI Reels translationA Reel you posted; lip sync is optionalThe languages Meta lists on its page
Sume Avatar 1.0 talking videoAn avatar plus a scriptEnglish only
VEED Fabric 1.0A still plus Sume-hosted audioSet by the audio
MiniMax H3 Max Lip SyncThe same still plus audio body, audio 5 to 14.8 sSet by the audio

How do I make a non-English talking shot?

Generate the voice first with POST /v1/tts-1.0/generate and set language to the language of the script; the OpenAPI says to set it for every non-English transcript because an omitted value defaults to English at the provider. Then send the resulting audio URL and your still to the Fabric endpoint.

The Sume docs do not claim an avatar's own voice speaks other languages, so listen to the output before you commit a batch.

Which should I choose?

If the video is already on Instagram or Facebook and the language is on Meta's list, use the platform option. If you need a talking shot in a language Avatar 1.0 does not speak, build it from TTS and a lip-sync endpoint. For the longer comparison of dubbing and lip sync, see lip sync vs dubbing.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume