Face swap, motion control or lip sync: which Sume avatar route?
Pick the Sume avatar route by what you already have: a clip with a person (face swap), a motion video (Kling motion control) or a voice line (H3 Max lip sync).
Choose by the asset you already hold. If you have a finished clip with a person in it and want your avatar in their place, use face swap. If you have a reference motion video and want a still to perform it, use Kling motion control. If you have a spoken line and want a still to say it, use H3 Max lip sync. If you only have a script, use the Avatar 1.0 talking-video route and let Sume voice it.
All four take an avatar from Sume, so the avatar is the constant and the input is the variable. The routes and limits below come from Models, Face swap, Generate avatar video and Create new avatar, read 2026-10-06.
What does each route need and cost?
| Route | You supply | Length | Rate |
|---|---|---|---|
| Face swap (Beta) | avatar_handle, video_url, quality | About 4 to 15 s source with audio | $0.184 (standard), $0.245 (plus) or $0.55 (max) per second |
| Kling 3.0 motion control | Still or avatar, motion_video_url | 1 to 30 s | $0.1575 per second |
| H3 Max lip sync | Still or avatar, Sume-hosted audio_url | 5 to 14.8 s | $0.0625, $0.10 or $0.20 per second |
| Avatar 1.0 talking video | avatar_handle, script or video_inputs | 4 to 60 s | $0.184, $0.245 or $0.55 per second by quality |
Where do people pick the wrong one?
Motion control does not make a still speak in your words. Its sound comes from the reference video when keep_original_sound is on, and the stored motion-control post explains why. If the point is to say a line, you want lip sync, which animates the mouth to audio you supply.
The face-swap beta worker uses the same Kling motion-control queue for its motion step, then puts the source audio back, so the performance you get is always the source clip's. Face swap is the mirror case. It replaces a face in video you already shot, so the performance, body and voice are the source's, not yours. It is Beta, and it needs usable audio in the source clip. If you want your avatar to say new words, swapping is the wrong tool; make a talking video instead.
Lip sync has the tightest audio window, 5 to 14.8 seconds, and only accepts audio hosted by Sume. For longer lines either split the voiceover or move to the talking-video route, which accepts up to 60 seconds.
A decision in four questions
Each route reserves cost when the job is admitted and refunds it if the job fails, so a wrong first choice costs you only the jobs that finish. Run the cheapest representative clip first. For the face-swap side in detail, the face swap beta post lists its limits and the recast comparison covers whole-body swaps.
- Do you already have footage of someone doing the thing? Face swap, if the clip is about 4 to 15 seconds.
- Do you have a motion you want copied onto a character? Motion control.
- Do you have audio and a face? H3 Max lip sync.
- Do you have only words? Avatar talking video.
How do you run a fair test across routes?
Use one avatar and one small brief, then run each route that applies. A 10-second line through lip sync at 768p costs $1.00; the same duration through Kling motion control costs $1.575. Compare them on whether the result serves the brief, not on which is more impressive. Write the finding next to the cost so a colleague can see why a route was chosen.
If two routes both fit, prefer the one whose failure is cheapest to retry. A lip-sync job with a bad audio file wastes only that job, and the avatar stays untouched. Whichever route you pick, store the avatar handle, input URLs and key with the output so the clip can be reproduced.
Sources
Related posts
More in Sume Avatar 1.0
- Pipecat avatar TTFB metrics vs timing a Sume avatar job
Pipecat's Tavus service reports TTFB from TTSStartedFrame and BotStartedSpeakingFrame. A Sume avatar job has no first byte to time: measure submit to completed.
- Pipecat HeyGen LiveAvatar: duration by tier vs a 60 second Sume job
In Pipecat, HeyGen LiveAvatar session length depends on your subscription tier. On Sume the length rule is fixed: 4 to 60 seconds per job, priced per second.
- Pipecat Simli max_session_length vs Sume's 4 to 60 second clip
Simli in Pipecat caps a live session with max_session_length and max_idle_time. Sume has no session: an avatar job is a 4 to 60 second file.
- Pipecat TavusVideoService live avatar, or a recorded Sume clip?
Pipecat's Tavus service speaks your agent's TTS live over WebRTC. A Sume avatar job is a 4 to 60 second recorded MP4. How to pick, by what the viewer does.
Written by Sume