Face swap, motion control or lip sync: which Sume avatar route?

Pick the Sume avatar route by what you already have: a clip with a person (face swap), a motion video (Kling motion control) or a voice line (H3 Max lip sync).

5 min readSume
All posts

Choose by the asset you already hold. If you have a finished clip with a person in it and want your avatar in their place, use face swap. If you have a reference motion video and want a still to perform it, use Kling motion control. If you have a spoken line and want a still to say it, use H3 Max lip sync. If you only have a script, use the Avatar 1.0 talking-video route and let Sume voice it.

All four take an avatar from Sume, so the avatar is the constant and the input is the variable. The routes and limits below come from Models, Face swap, Generate avatar video and Create new avatar, read 2026-10-06.

What does each route need and cost?

Inputs and limits from docs.sume.com, read 2026-10-06. Rates are the Sume price list x 1.25 for Kling and H3 Max lip sync; the face-swap rate is the Avatar Video rate by quality.
RouteYou supplyLengthRate
Face swap (Beta)avatar_handle, video_url, qualityAbout 4 to 15 s source with audio$0.184 (standard), $0.245 (plus) or $0.55 (max) per second
Kling 3.0 motion controlStill or avatar, motion_video_url1 to 30 s$0.1575 per second
H3 Max lip syncStill or avatar, Sume-hosted audio_url5 to 14.8 s$0.0625, $0.10 or $0.20 per second
Avatar 1.0 talking videoavatar_handle, script or video_inputs4 to 60 s$0.184, $0.245 or $0.55 per second by quality

Where do people pick the wrong one?

Motion control does not make a still speak in your words. Its sound comes from the reference video when keep_original_sound is on, and the stored motion-control post explains why. If the point is to say a line, you want lip sync, which animates the mouth to audio you supply.

The face-swap beta worker uses the same Kling motion-control queue for its motion step, then puts the source audio back, so the performance you get is always the source clip's. Face swap is the mirror case. It replaces a face in video you already shot, so the performance, body and voice are the source's, not yours. It is Beta, and it needs usable audio in the source clip. If you want your avatar to say new words, swapping is the wrong tool; make a talking video instead.

Lip sync has the tightest audio window, 5 to 14.8 seconds, and only accepts audio hosted by Sume. For longer lines either split the voiceover or move to the talking-video route, which accepts up to 60 seconds.

A decision in four questions

Each route reserves cost when the job is admitted and refunds it if the job fails, so a wrong first choice costs you only the jobs that finish. Run the cheapest representative clip first. For the face-swap side in detail, the face swap beta post lists its limits and the recast comparison covers whole-body swaps.

  • Do you already have footage of someone doing the thing? Face swap, if the clip is about 4 to 15 seconds.
  • Do you have a motion you want copied onto a character? Motion control.
  • Do you have audio and a face? H3 Max lip sync.
  • Do you have only words? Avatar talking video.

How do you run a fair test across routes?

Use one avatar and one small brief, then run each route that applies. A 10-second line through lip sync at 768p costs $1.00; the same duration through Kling motion control costs $1.575. Compare them on whether the result serves the brief, not on which is more impressive. Write the finding next to the cost so a colleague can see why a route was chosen.

If two routes both fit, prefer the one whose failure is cheapest to retry. A lip-sync job with a bad audio file wastes only that job, and the avatar stays untouched. Whichever route you pick, store the avatar handle, input URLs and key with the output so the clip can be reproduced.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume