Face swap or talking video: which Avatar endpoint fits?
Pick Sume Avatar face swap (Beta) or Avatar talking video by what you have: footage with audio, or a script. Inputs, limits and polling side by side.
Use Sume Avatar face swap (Beta) when you already have a short source video with usable audio and only the face should change. Use the Avatar talking-video endpoint when you have a script and want the avatar to speak it. The two take different inputs, and face swap accepts no prompt, transcript or aspect ratio.
The comparison below comes from Face swap (Beta) and Generate avatar video. Both need a ready avatar handle.
Side by side
Both endpoints are job-backed, so the polling and result routes are the same. What differs is the input contract and what you can steer.
| Question | Face swap (Beta) | Talking video |
|---|---|---|
| Route | POST /v1/models/sume/avatar-face-swap/v1.0/runs | POST /v1/avatar-1.0/talking-video |
| You provide | avatar_handle, public HTTPS video_url, quality | avatar_handle, script or video_inputs |
| Quality default | None; quality is required | plus when omitted |
| Length | Source about 4-15 seconds with usable audio | Estimated 4-60 seconds |
| Prompts, transcripts | Not supported | Script is the input |
| Aspect ratio | Not supported (the source decides) | 1:1, 3:4, 9:16, 4:3, 16:9 |
| Captions | Not in the contract | Optional inline captions |
Source footage rules for face swap
The video_url must be a fetchable public HTTPS video. The endpoint rejects localhost, private-network, non-HTTPS and signed or private URLs, and provider task URLs. In Beta, the docs describe the current plan as approximately 4-15 seconds with usable audio. A silent or very long source is outside that plan.
Because there is no default, a request without quality is refused rather than silently priced at a middle tier. Choose standard, plus or max per call.
When the script is the point
Pick talking video when the words are yours. It supports single scripts or ordered scenes, silence beats, an optional product image, a prompted or photo scene, and optional inline captions. It also has a preview stage: you can approve first-frame stills before the full render.
Talking video is English-only for speech, so a face swap can be the more practical route when your source footage is in another language. That is a judgment call from the English-only limit, not a documented swap feature, so test one clip.
Handling the result
Poll GET /v1/jobs/:id/status and then /result, or use a webhook, as you would for any generation. Use resource_status as the main readiness signal and job_status while polling. A finished resource exposes a public video under media.sume.com.
To check a candidate clip before spending, run it through reference ingest first with purpose: face_swap where that tool is enabled. Sume stores the purpose but does not interpret it, so you read the shots and audio facts yourself.
Questions to ask before you choose
Start from your assets, not from the feature list. If your footage already has the performance you want (timing, gestures, a real delivery) and the only change is who appears, face swap keeps all of that and replaces the identity. If you have no footage, there is nothing to swap, and talking video is the only fit.
Think about length next. A face-swap source is planned for roughly 4-15 seconds in Beta, while talking video accepts up to 60 seconds of estimated speech, so a longer piece needs talking video or several swaps stitched in an edit. Then think about review: talking video offers a first-frame preview before the full render, while face swap goes straight to a job, so test with your shortest clip first.
Last, check the avatar. Both routes need a ready avatar handle, created once and reused. Both docs pages say the avatar must be ready, so confirm it is before you build a pipeline on it.
Sources
Related posts
More in Sume Avatar 1.0
- Gradium's 329 character voices vs voices of Sume avatars
Gradium lists 329 character voices. On Sume, a spoken character is an avatar with a ready voice; Avatar 1.0 is English-only. What that means for casting.
- Moving Company Spokesperson Ad: A 20-Second Avatar for About $4.90
A 20-second plus-quality Avatar 1.0 spokesperson clip costs about $4.90 on Sume, at 98 cents per 4 seconds. English only, with captions at $0.20 extra.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume