Face swap or avatar talking video: which Sume endpoint?
Use face swap when you already have a 4-15 second source video; use the avatar talking video when you have a script. Inputs, limits and what each returns.
Pick Beta face swap when a source video already exists and you want a ready avatar's face applied to it; pick the avatar talking video when you have a script and want the avatar to perform it. Both need an avatar first, and the face swap docs say it is not for script-driven talking video.
What does each endpoint take?
Both use a ready avatar referenced by avatar_handle. Face swap adds a public HTTPS video_url and a required quality. Avatar video adds exactly one of script or video_inputs, with optional product image, scene, quality, aspect_ratio and captions.
| Face swap (Beta) | Avatar talking video | |
|---|---|---|
| Route | POST /v1/models/sume/avatar-face-swap/v1.0/runs | POST /v1/avatar-1.0/talking-video |
| Main input | A public source video | A script or scene plan |
| Length | About 4-15 s source | 4-60 s estimated |
| Quality | Required | Optional, default plus |
| Aspect ratio | Not supported | 1:1, 3:4, 9:16, 4:3, 16:9 |
| Captions | Not documented | Optional inline captions |
When is face swap the right one?
When you have footage: a short clip with usable audio, with the avatar's face on it. The docs also say it is not the old consumer-product /face-swap route; use only the Developer API path above.
When is the talking video the right one?
When you start from words. Avatar video reads a script, supports multi-scene plans with silence beats, and can burn captions in. Requests outside 4-60 estimated seconds are rejected, so long scripts must be split. Details in the avatar video docs.
Do I need to make the avatar in both cases?
Yes. Create it once from a prompt, traits or photo in the avatar docs and reuse the handle. Getting consent for any real person's face is on you, since the docs show no consent step.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen 30-minute avatar video: what Sume does in one request
HeyGen says it can make a 30-minute talking video in one pass. One Sume avatar video request covers 4-60 seconds, so longer pieces are split into jobs.
- HeyGen Avatar V's 15-second recording vs a Sume photo avatar
HeyGen's Avatar V starts from a 15-second recording. Sume's avatar API starts from a prompt, traits or one public HTTPS image, with no recording step.
- HeyGen Edit Look: retouching an AI avatar, and the Sume route
HeyGen's Edit Look retouches an existing avatar in place. Sume's docs show no such edit, so the route is a new avatar from a retouched photo and a new handle.
- Shortest video an AI dub or face swap accepts: 5 s vs 4 s
Synthesia's dubbing page says a video must be at least 5 seconds. Sume's Beta face swap plans for about 4-15 seconds and avatar videos take 4-60. Side by side.
Written by Sume