Face swap or talking video: which Avatar endpoint fits?

Pick Sume Avatar face swap (Beta) or Avatar talking video by what you have: footage with audio, or a script. Inputs, limits and polling side by side.

5 min readSume
All posts

Use Sume Avatar face swap (Beta) when you already have a short source video with usable audio and only the face should change. Use the Avatar talking-video endpoint when you have a script and want the avatar to speak it. The two take different inputs, and face swap accepts no prompt, transcript or aspect ratio.

The comparison below comes from Face swap (Beta) and Generate avatar video. Both need a ready avatar handle.

Side by side

Both endpoints are job-backed, so the polling and result routes are the same. What differs is the input contract and what you can steer.

Face swap (Beta) against Avatar talking video (docs read 2026-10-10)
QuestionFace swap (Beta)Talking video
RoutePOST /v1/models/sume/avatar-face-swap/v1.0/runsPOST /v1/avatar-1.0/talking-video
You provideavatar_handle, public HTTPS video_url, qualityavatar_handle, script or video_inputs
Quality defaultNone; quality is requiredplus when omitted
LengthSource about 4-15 seconds with usable audioEstimated 4-60 seconds
Prompts, transcriptsNot supportedScript is the input
Aspect ratioNot supported (the source decides)1:1, 3:4, 9:16, 4:3, 16:9
CaptionsNot in the contractOptional inline captions

Source footage rules for face swap

The video_url must be a fetchable public HTTPS video. The endpoint rejects localhost, private-network, non-HTTPS and signed or private URLs, and provider task URLs. In Beta, the docs describe the current plan as approximately 4-15 seconds with usable audio. A silent or very long source is outside that plan.

Because there is no default, a request without quality is refused rather than silently priced at a middle tier. Choose standard, plus or max per call.

When the script is the point

Pick talking video when the words are yours. It supports single scripts or ordered scenes, silence beats, an optional product image, a prompted or photo scene, and optional inline captions. It also has a preview stage: you can approve first-frame stills before the full render.

Talking video is English-only for speech, so a face swap can be the more practical route when your source footage is in another language. That is a judgment call from the English-only limit, not a documented swap feature, so test one clip.

Handling the result

Poll GET /v1/jobs/:id/status and then /result, or use a webhook, as you would for any generation. Use resource_status as the main readiness signal and job_status while polling. A finished resource exposes a public video under media.sume.com.

To check a candidate clip before spending, run it through reference ingest first with purpose: face_swap where that tool is enabled. Sume stores the purpose but does not interpret it, so you read the shots and audio facts yourself.

Questions to ask before you choose

Start from your assets, not from the feature list. If your footage already has the performance you want (timing, gestures, a real delivery) and the only change is who appears, face swap keeps all of that and replaces the identity. If you have no footage, there is nothing to swap, and talking video is the only fit.

Think about length next. A face-swap source is planned for roughly 4-15 seconds in Beta, while talking video accepts up to 60 seconds of estimated speech, so a longer piece needs talking video or several swaps stitched in an edit. Then think about review: talking video offers a first-frame preview before the full render, while face swap goes straight to a job, so test with your shortest clip first.

Last, check the avatar. Both routes need a ready avatar handle, created once and reused. Both docs pages say the avatar must be ready, so confirm it is before you build a pipeline on it.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume