Face swap beta: motion stage runs on Kling, source audio muxed back

Sume's Avatar Face Swap beta runs its motion stage on the Kling 3.0 motion control queue, strips the source audio, then muxes the original audio back in.

5 min readSume
All posts

Sume's Avatar Face Swap beta does its motion step on the same Kling 3.0 Standard motion control queue that the public motion control route uses. The swapped first frame is the character, the source video (with its audio removed) is the motion, and the finalize stage muxes the original audio back in. So the sound in your result is the sound of your source, not something the provider made.

The stages in plain words

A face swap job is a small pipeline, and the motion step is one stage of it. The pipeline swaps the face on the first frame of your video, then uses that frame as the still for a Kling motion pass that follows the movement of the source clip, then joins the original audio to the result.

The audio is taken out for the motion stage because the motion model is not asked to produce a soundtrack for a swap. The source audio is the truth, and the finalize stage puts it back so that lip movement, music and room tone stay as they were.

What that means for a buyer

Three practical points follow.

  • Length and cost are tied to the source clip, as they are for any motion pass. The beta accepts short clips, so read the current limits on the face swap page before you plan a batch.
  • The soundtrack of the output is your own. If the source has music that you do not have rights to, the swapped clip has it too.
  • Because it shares the queue, a busy motion control backlog can slow a face swap, and the reverse is also true. Plan concurrency across both.

Face swap and motion control side by side

Avatar Face Swap beta and public Kling motion control on Sume (read 2026-10-05)
ItemFace swap betaMotion control route
Provider queueKling 3.0 Standard motion controlKling 3.0 Standard motion control
Character stillswapped first frame of your videoimage_url or avatar you send
Motion sourceyour video, audio removedmotion_video_url you send
Audio in resultoriginal source audio, muxed backkeep_original_sound (default true)
How you call itface swap beta endpointPOST /v1/kling/3.0/motion-control

When to call which

Call the face swap beta when your input is a finished video and you want the person in it replaced by a Sume avatar. You hand over one video and one avatar, and the pipeline handles the first frame and the motion.

Call motion control directly when you already hold a still and a performance clip, and you control both. You get the length, the sound switch and the orientation setting yourself, and a clean price of $0.1575 a second.

Both end in a normal Sume job, so the completion path is the same. Use a webhook instead of a poll loop, verify the signature, and then fetch the result. The face swap webhook post covers the event name and the headers for that case.

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume