Does Sume face swap keep the original voice, outfit and scene?
Sume's Face Swap (Beta) keeps the source clip's audio, camera and timing but replaces the whole person, not just the face. What stays, what changes, cost.

Yes. In Sume's Face Swap (Beta) the source clip keeps its own audio: whatever the original person said, in their voice, is what the finished video plays. The picture changes more than the name suggests, though. Sume's workflow is written to replace the whole person with your avatar, face, hair, outfit, hands and upper body, while holding the source's camera, scene, timing and length.
That makes Face Swap a different tool from the avatar talking video. A talking video speaks a script in the avatar's voice. A face swap re-casts a performance you already have. This page lists what carries over from the source, what does not, and how to confirm the audio on your own result. Everything below is read from Sume's Face swap (Beta) docs and the API reference on 2026-10-03, plus Sume's workflow code for the stages.
What stays from the source clip, and what is replaced?
The Beta endpoint takes an avatar_handle, a public HTTPS video_url and a required quality, and nothing else. The stages behind it work from the source's own frame and sound, so most of what you filmed survives.
| Part of the clip | Result | Where it comes from |
|---|---|---|
| Spoken audio | Source audio, not the avatar's voice | Workflow muxes the source audio back after the video stage |
| Camera framing and timing | Follows the source | Source clip drives the motion |
| Scene, lighting, product or prop | Follows the source | Opening frame is built from a source frame |
| Clip length | Follows the source (about 4-15 s) | No duration field exists on the request |
| Face, hair, outfit, hands | The avatar's | Avatar reference is the identity target |
| Language of the speech | Whatever the source speaks | There is no language field on this endpoint |
Why does the avatar not speak in its own voice?
Because no voice is generated. The request has no script, transcript or voice field, and Sume's reference says prompts, transcripts, duration knobs, aspect ratio and generated-audio flags are intentionally unsupported. The source speech is transcribed internally only as guidance for mouth shapes, then the original audio is re-attached to the finished video.
If you need the avatar to say new words in its voice, that is the talking-video route with a script, not a face swap. If you need both, a performance from your footage and new words, you are choosing between two clips, not one.
What does that mean for a brand clip?
- The person in the source must be someone you can re-cast: a founder, an employee, or a creator who signed a release that covers an AI replacement.
- A swap is only as clean as the source audio. The Beta expects usable speech; a silent source is out of contract.
- Anything said in the source is still said. Claims in the original take become claims in your ad, with your avatar attached.
- The result is a new person saying real words, so label and disclose it the way you would any avatar ad.
How do I confirm the result kept the audio?
Fetch the finished job's video from its media.sume.com URL and run video inspect on it. The probe and stills are free; adding transcribe: true reserves the speech-to-text per-minute rate and returns word timings, so you can compare the transcript with the source. If the result has no audio stream, the probe says so before you publish.
Poll resource_status for readiness and job_status for the job, as the docs advise. A completed job whose video_url is empty is a resource-state question, not a lost render.
What should I film to get a good swap?
Because the swap follows the source, the source decides most of the result. The workflow is written to keep the camera locked to what you filmed, so a steady phone shot, one person facing the lens and speaking clearly gives it the least to guess. Sume's instructions to the model call for matching the source's lighting direction and colour, and for keeping any product or prop where it was, so a prop held in the source stays in the swapped clip.
Keep the clip between about 4 and 15 seconds. The Beta contract says worker validation currently targets that range with usable audio. A silent clip fails, and a clip far past 15 seconds is outside what the endpoint reserves for, so trim first and swap the clip you intend to publish. For a longer piece, swap several short takes and join them on a timeline.
Avoid sources where the person is partly hidden, turned away or one of several people on screen. Sume's docs do not describe multi-person handling for Face Swap, so treat that as unsupported and try a short test clip before a batch.
Can I use this for a localisation or voice change?
No. Face Swap has no language knob and no voice field, and its speech stage only listens to the source. A swap cannot translate the clip or give it your avatar's voice. If the source speaks Spanish, the result speaks Spanish with an English-speaking avatar's face, because the audio is simply carried across. Avatar 1.0 itself is English-only in code for scripts, so do not read a swap as a path to non-English avatar speech with the avatar's own voice; for dubbing there is a separate route covered in dubbing an avatar video with lip sync.
What does it cost?
Face Swap (Beta) reserves a 15-second maximum at the Avatar Video rate for the tier you pass, with no product image: $2.76 on standard, $3.68 on plus and $8.25 on max, per Sume's pricing code. There is no default tier, so omit quality and the request is rejected. Creating the avatar itself is a separate one-time $0.95. The full ladder is in Face swap cost ceiling.
Sources
Related posts
More in Sume Avatar 1.0
- How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
- Shortest AI avatar video: 4 seconds, from $0.74 on Sume Standard
Sume avatar videos run 4 to 60 seconds. Per-second rates for Standard, Plus and Max, the 4 second floor, and what 15, 30 and 60 second clips cost.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
Written by Sume