Does Sume face swap keep the original voice, outfit and scene?

Sume's Face Swap (Beta) keeps the source clip's audio, camera and timing but replaces the whole person, not just the face. What stays, what changes, cost.

5 min readSume
All posts

Yes. In Sume's Face Swap (Beta) the source clip keeps its own audio: whatever the original person said, in their voice, is what the finished video plays. The picture changes more than the name suggests, though. Sume's workflow is written to replace the whole person with your avatar, face, hair, outfit, hands and upper body, while holding the source's camera, scene, timing and length.

That makes Face Swap a different tool from the avatar talking video. A talking video speaks a script in the avatar's voice. A face swap re-casts a performance you already have. This page lists what carries over from the source, what does not, and how to confirm the audio on your own result. Everything below is read from Sume's Face swap (Beta) docs and the API reference on 2026-10-03, plus Sume's workflow code for the stages.

What stays from the source clip, and what is replaced?

The Beta endpoint takes an avatar_handle, a public HTTPS video_url and a required quality, and nothing else. The stages behind it work from the source's own frame and sound, so most of what you filmed survives.

Face Swap (Beta): source clip versus result, read 2026-10-03
Part of the clipResultWhere it comes from
Spoken audioSource audio, not the avatar's voiceWorkflow muxes the source audio back after the video stage
Camera framing and timingFollows the sourceSource clip drives the motion
Scene, lighting, product or propFollows the sourceOpening frame is built from a source frame
Clip lengthFollows the source (about 4-15 s)No duration field exists on the request
Face, hair, outfit, handsThe avatar'sAvatar reference is the identity target
Language of the speechWhatever the source speaksThere is no language field on this endpoint

Why does the avatar not speak in its own voice?

Because no voice is generated. The request has no script, transcript or voice field, and Sume's reference says prompts, transcripts, duration knobs, aspect ratio and generated-audio flags are intentionally unsupported. The source speech is transcribed internally only as guidance for mouth shapes, then the original audio is re-attached to the finished video.

If you need the avatar to say new words in its voice, that is the talking-video route with a script, not a face swap. If you need both, a performance from your footage and new words, you are choosing between two clips, not one.

What does that mean for a brand clip?

  • The person in the source must be someone you can re-cast: a founder, an employee, or a creator who signed a release that covers an AI replacement.
  • A swap is only as clean as the source audio. The Beta expects usable speech; a silent source is out of contract.
  • Anything said in the source is still said. Claims in the original take become claims in your ad, with your avatar attached.
  • The result is a new person saying real words, so label and disclose it the way you would any avatar ad.

How do I confirm the result kept the audio?

Fetch the finished job's video from its media.sume.com URL and run video inspect on it. The probe and stills are free; adding transcribe: true reserves the speech-to-text per-minute rate and returns word timings, so you can compare the transcript with the source. If the result has no audio stream, the probe says so before you publish.

Poll resource_status for readiness and job_status for the job, as the docs advise. A completed job whose video_url is empty is a resource-state question, not a lost render.

What should I film to get a good swap?

Because the swap follows the source, the source decides most of the result. The workflow is written to keep the camera locked to what you filmed, so a steady phone shot, one person facing the lens and speaking clearly gives it the least to guess. Sume's instructions to the model call for matching the source's lighting direction and colour, and for keeping any product or prop where it was, so a prop held in the source stays in the swapped clip.

Keep the clip between about 4 and 15 seconds. The Beta contract says worker validation currently targets that range with usable audio. A silent clip fails, and a clip far past 15 seconds is outside what the endpoint reserves for, so trim first and swap the clip you intend to publish. For a longer piece, swap several short takes and join them on a timeline.

Avoid sources where the person is partly hidden, turned away or one of several people on screen. Sume's docs do not describe multi-person handling for Face Swap, so treat that as unsupported and try a short test clip before a batch.

Can I use this for a localisation or voice change?

No. Face Swap has no language knob and no voice field, and its speech stage only listens to the source. A swap cannot translate the clip or give it your avatar's voice. If the source speaks Spanish, the result speaks Spanish with an English-speaking avatar's face, because the audio is simply carried across. Avatar 1.0 itself is English-only in code for scripts, so do not read a swap as a path to non-English avatar speech with the avatar's own voice; for dubbing there is a separate route covered in dubbing an avatar video with lip sync.

What does it cost?

Face Swap (Beta) reserves a 15-second maximum at the Avatar Video rate for the tier you pass, with no product image: $2.76 on standard, $3.68 on plus and $8.25 on max, per Sume's pricing code. There is no default tier, so omit quality and the request is rejected. Creating the avatar itself is a separate one-time $0.95. The full ladder is in Face swap cost ceiling.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume