Avatar face swap or talking video: recorded presenter vs scripted one

Sume has two ways to put an avatar on screen: face swap onto a 4 to 15 second source clip you recorded, or a script-driven talking video. How to choose.

5 min readSume
All posts

Choose face swap when someone has already recorded the performance and you only want the avatar's face on it; choose the talking video when you have a script and no footage. Both use a ready Sume avatar handle, but face swap is a Beta endpoint with a source clip of about 4 to 15 seconds and usable audio, and talking video is the production route with a 4 to 60 second window.

Face swap vs talking video, per the Sume docs as of 2026-10-08
QuestionAvatar face swap (Beta)Avatar talking video
Inputavatar_handle, public HTTPS video_url, qualityavatar_handle, script or video_inputs
Who performsA person in your source clipThe model, from your script
LengthSource about 4 to 15 s with usable audio4 to 60 s estimated
Prompts and aspect ratioNot supported by designaspect_ratio 1:1, 3:4, 9:16, 4:3, 16:9
Quality fieldRequired, no defaultOptional, default plus

When the performance matters

A face swap keeps the timing, gestures and voice of the source clip. If a founder, creator or actor records a take with the right energy, the swap changes who appears on screen without rewriting the performance. The endpoint rejects prompts, transcripts and duration knobs, so you cannot steer it with text. The source must be a public HTTPS URL that is not signed or private.

The performer's voice stays in the output, since the audio of your clip is the audio of the result. If you do not own the right to that voice, do not use the clip.

When the script matters

A talking video is the better choice when many variants share a message or the words change often. You edit a script, review first-frame stills with an Avatar video preview, then render. You can add a product image or a scene reference, and captions can be burned in at generation time.

Talking video speaks English only in the pipeline prompt, so a recorded non-English performance is a case where face swap is the only avatar route.

Cost shape

Talking video bills per second at $0.184, $0.245 or $0.55 for standard, plus and max. The face swap catalog entry is a Beta contract estimate: the standard talking-video rate times the 15 second source maximum, so 15 x 0.184 = $2.76 on standard, 15 x 0.245 = $3.675 on plus and 15 x 0.55 = $8.25 on max. Read the actual amount from your job result.

Both are asynchronous jobs. Neither streams, and neither reacts to a person talking to it.

A note on practice

Whichever route you pick, get permission from every person whose likeness or voice appears, in writing. A face swap puts an avatar identity on a real performer's recording, so the performer's consent matters as much as the avatar owner's. Keep both records with the job id.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume