Perso AI dubbing: 10 speakers, 2-speaker lip sync, vs Sume

Perso says it detects up to 10 speakers and lip-syncs two. Sume's avatar video uses one avatar per final video. What to use for a multi-speaker dub.

5 min readSume
All posts

Perso AI's dubbing page claims it auto-detects up to 10 speakers per scene, includes 2-speaker lip sync, keeps a 98% voice match and covers 99+ languages (Perso, read 2026-10-02). Sume cannot do that in one call: its avatar video renders one avatar per final video, and there is no speaker-labelled dubbing route.

What Perso claims

The page also lists an editable script for technical terms, an API at its developer site, and one free dub on signup. These are marketing claims from the vendor; the page does not give a file length limit. Treat the voice-match figure as Perso's own claim, not a measured number.

Perso AI dubbing page, read 2026-10-02
ItemClaim on the page
Languages99+
SpeakersAuto-detects up to 10 per scene
Lip syncIncluded, with 2-speaker lip sync
Script editingEditable script for technical terms
APIA Dubbing API, at the developer site
File length limitNot specified on the page

What Sume has for speakers

The avatar video docs say current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. Multi-scene video_inputs can include silence beats, but they do not put two speaking faces in one frame. Sume's speech to text request has no speaker-label field.

A two-speaker workaround on Sume

This takes manual work, but each step is explicit.

  • Transcribe each speaker's isolated audio file separately if you have them.
  • Speak each translated line with its own library voice.
  • Lip sync each speaker's face as its own clip, using a still image plus 5 to 14.8 seconds of audio.
  • Join the clips on a Timeline and join the audio with Timeline audio.

Honest limits

The lip-sync route needs a still image and audio between 5 and 14.8 seconds, so it does not re-sync mouths in existing footage. If your content is two on-camera speakers in a single shot, a dedicated dubbing tool such as Perso fits better than a Sume chain. The lip sync vs dubbing post explains the difference.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume