Face swap or avatar video? Pick by whether you already have footage

Sume's face swap Beta applies an avatar face to your own 4-15 second video; Avatar Video renders a new clip from a script, up to 60 seconds. How to choose.

5 min readSume
All posts

Short answer

Use face swap if you already have a short video you want to keep and only the face should change. Use Avatar Video if you have words and need a new clip made around a presenter. Sume's docs draw this line themselves: face swap is for people who have a ready avatar and a short public source video, and they point to Avatar Video when you need script-driven generation. The two share a rate card but differ in input, length and control.

The two side by side

Both start from a ready Avatar 1.0 handle. After that, the inputs and limits differ.

Face swap Beta versus Avatar Video (as of 2026-10-08)
PointFace swap (Beta)Avatar Video
InputPublic HTTPS source video plus avatar_handleScript or video_inputs plus avatar_handle
LengthAbout 4 to 15 seconds, usable audio4 to 60 seconds, from the script
qualityRequired, no defaultstandard, plus (default) or max
Prompts and transcriptsNot acceptedScript and scene prompt accepted
Aspect ratioNot a parameter1:1, 3:4, 9:16, 4:3 or 16:9
RoutePOST /v1/models/sume/avatar-face-swap/v1.0/runsPOST /v1/avatar-1.0/talking-video

Cost for a 15-second result

Fifteen seconds is the face-swap ceiling and a normal Avatar Video length. At the same tier, the per-second figure is the same, so the cost for 15 seconds is the same too. Face swap prices its largest estimate with the Avatar Video no-product rate times the 15-second maximum.

15 seconds on either route (rates as of 2026-10-08)
TierRate per second15 seconds
standard$0.184$2.76
plus$0.245$3.68
max$0.55$8.25

Decision rules

Choose by the asset you hold. A recorded product demo, a clip of an actor, or a vertical take that works as footage points to face swap. A script, a talking point list or a launch note points to Avatar Video. If the source clip is longer than about 15 seconds, face swap will not take it as is; trim it, or switch to a script.

  • Have footage, 4 to 15 seconds, with audio: face swap.
  • Have words, 4 to 60 seconds: Avatar Video.
  • Need captions burned in: Avatar Video has an inline captions field.
  • Need a first-frame check: Avatar Video has a preview route.

Disclosure applies to both

Either route can produce a realistic synthetic person. The YouTube help page asks creators to disclose realistic altered or synthetic content, including content that makes a real person appear to say or do something they did not. Face swap onto footage of someone else sits closest to that line. In both cases, you set the label at upload and keep a record of who approved the result.

Practical sequencing helps. Create the avatar once ($0.95) and reuse the handle across both routes. Try the face swap on the shortest clip that shows the problem, such as 4 seconds at standard for 4 x $0.184 = $0.74, before sending a 15-second clip. The Beta requires quality on every request, so pick it deliberately: there is no default, and a request without it is not accepted. Both routes return a job, so the polling code you write for one carries over to the other, and both can use async, sync, or webhook modes.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume