Perso AI dubbing: 10 speakers, 2-speaker lip sync, vs Sume
Perso says it detects up to 10 speakers and lip-syncs two. Sume's avatar video uses one avatar per final video. What to use for a multi-speaker dub.
Perso AI's dubbing page claims it auto-detects up to 10 speakers per scene, includes 2-speaker lip sync, keeps a 98% voice match and covers 99+ languages (Perso, read 2026-10-02). Sume cannot do that in one call: its avatar video renders one avatar per final video, and there is no speaker-labelled dubbing route.
What Perso claims
The page also lists an editable script for technical terms, an API at its developer site, and one free dub on signup. These are marketing claims from the vendor; the page does not give a file length limit. Treat the voice-match figure as Perso's own claim, not a measured number.
| Item | Claim on the page |
|---|---|
| Languages | 99+ |
| Speakers | Auto-detects up to 10 per scene |
| Lip sync | Included, with 2-speaker lip sync |
| Script editing | Editable script for technical terms |
| API | A Dubbing API, at the developer site |
| File length limit | Not specified on the page |
What Sume has for speakers
The avatar video docs say current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene. Multi-scene video_inputs can include silence beats, but they do not put two speaking faces in one frame. Sume's speech to text request has no speaker-label field.
A two-speaker workaround on Sume
This takes manual work, but each step is explicit.
- Transcribe each speaker's isolated audio file separately if you have them.
- Speak each translated line with its own library voice.
- Lip sync each speaker's face as its own clip, using a still image plus 5 to 14.8 seconds of audio.
- Join the clips on a Timeline and join the audio with Timeline audio.
Honest limits
The lip-sync route needs a still image and audio between 5 and 14.8 seconds, so it does not re-sync mouths in existing footage. If your content is two on-camera speakers in a single shot, a dedicated dubbing tool such as Perso fits better than a Sume chain. The lip sync vs dubbing post explains the difference.
Sources
Related posts
More in Comparisons
- Pexels free stock footage vs generated B-roll: what to use when
Compare Pexels licensed footage with generated B-roll: licence limits, control, cost and where each wins, with Pexels terms to re-check.
- Pictory video minutes per dollar vs Sume timeline render
Pictory lists 200 to 1,800 video minutes a month at $0.066 to $0.145 per minute. Sume renders a timeline at $0.10 per output minute. Dated 2026-10-01.
- Pika 2.5 and Pikaframes: Pika's video menu vs Sume's catalog ids
Pika lists six video models, including its own Pika 2.5 and Pikaframes. Sume has ids for most of the others but none named Pika, so check the catalog first.
- PixVerse C1 API price per second, 360p to 1080p, and Sume options
PixVerse C1 on fal costs $0.03 to $0.12 a second by resolution and audio, up to 15 seconds at 1080p. The price table and which Sume models cover 15 seconds.
Written by Sume