Avatar face swap or talking video: recorded presenter vs scripted one
Sume has two ways to put an avatar on screen: face swap onto a 4 to 15 second source clip you recorded, or a script-driven talking video. How to choose.
Choose face swap when someone has already recorded the performance and you only want the avatar's face on it; choose the talking video when you have a script and no footage. Both use a ready Sume avatar handle, but face swap is a Beta endpoint with a source clip of about 4 to 15 seconds and usable audio, and talking video is the production route with a 4 to 60 second window.
| Question | Avatar face swap (Beta) | Avatar talking video |
|---|---|---|
| Input | avatar_handle, public HTTPS video_url, quality | avatar_handle, script or video_inputs |
| Who performs | A person in your source clip | The model, from your script |
| Length | Source about 4 to 15 s with usable audio | 4 to 60 s estimated |
| Prompts and aspect ratio | Not supported by design | aspect_ratio 1:1, 3:4, 9:16, 4:3, 16:9 |
| Quality field | Required, no default | Optional, default plus |
When the performance matters
A face swap keeps the timing, gestures and voice of the source clip. If a founder, creator or actor records a take with the right energy, the swap changes who appears on screen without rewriting the performance. The endpoint rejects prompts, transcripts and duration knobs, so you cannot steer it with text. The source must be a public HTTPS URL that is not signed or private.
The performer's voice stays in the output, since the audio of your clip is the audio of the result. If you do not own the right to that voice, do not use the clip.
When the script matters
A talking video is the better choice when many variants share a message or the words change often. You edit a script, review first-frame stills with an Avatar video preview, then render. You can add a product image or a scene reference, and captions can be burned in at generation time.
Talking video speaks English only in the pipeline prompt, so a recorded non-English performance is a case where face swap is the only avatar route.
Cost shape
Talking video bills per second at $0.184, $0.245 or $0.55 for standard, plus and max. The face swap catalog entry is a Beta contract estimate: the standard talking-video rate times the 15 second source maximum, so 15 x 0.184 = $2.76 on standard, 15 x 0.245 = $3.675 on plus and 15 x 0.55 = $8.25 on max. Read the actual amount from your job result.
Both are asynchronous jobs. Neither streams, and neither reacts to a person talking to it.
A note on practice
Whichever route you pick, get permission from every person whose likeness or voice appears, in writing. A face swap puts an avatar identity on a real performer's recording, so the performer's consent matters as much as the avatar owner's. Keep both records with the job id.
Sources
Related posts
More in Comparisons
- Six stills and a logo clip as references: Wan 3.0, Omni, H3 or Kling?
For a brand-kit video with six stills and one logo clip, Wan 3.0 and MiniMax H3 fit; Omni caps clips at 3 s; Kling 3 takes no references. 8 s prices inside.
- ChatGPT Image 2.5 low vs Ideogram 4.5 low: 1 cent vs 4 cents
Low tiers on Sume: ChatGPT Image 2.5 is 1 cent, Ideogram 4.5 is 4. Omit quality and GPT runs at high (7 cents), Ideogram at medium (8).
- ChatGPT Image 2 or 2.5 at high quality: 27 cents vs 7 on Sume
Sume lists ChatGPT Image 2 at a high 1024 rate of $0.211 (27 cents billable) and Image 2.5 at $0.0527 (7 cents). The docs describe xhigh and max for 2.5.
- Cheapest plan with a commercial license: Pika, ElevenLabs, Suno, Luma
ElevenLabs Starter is $6, Suno Pro $8, Luma Plus $30 and Pika Creator $35 on their pages. Free tiers lack commercial use. Sume bills per job.
Written by Sume