What happens to the audio in each Sume video-to-video tool
Recast and Avatar Face Swap keep the source audio, Kling motion control keeps it by default, Omni always produces native audio, Genjutsu takes no audio field.

The audio rule differs by tool, and getting it wrong means a clip that sounds like the wrong person. h3-max-recast keeps the source audio, the Beta Avatar Face Swap muxes the original source audio back, Kling 3.0 motion control keeps the driving video's sound unless you set keep_original_sound to false, gemini-omni-flash-1.1 always generates native synced audio, and higgsfield-genjutsu has no audio field at all (Video Router docs, Videos docs and Face swap docs, read 2026-10-03). Read the table below before you plan a voice.
In every tool here the new person's photo contributes a face, not a voice.
The audio behavior table
Each row is stated in the Sume docs; where a field is absent, the row says so rather than guessing.
| Tool | Audio behavior | Control you have |
|---|---|---|
h3-max-recast | Source audio kept in the output | None; plan the voice upstream |
| Avatar Face Swap (Beta) | Source needs usable audio; original audio is muxed back | None; prompts and provider fields unsupported |
| Kling 3.0 motion control | Driving video's sound kept by default | keep_original_sound: true by default, false gives a silent clip |
gemini-omni-flash-1.1 | Native synced audio always on; generate_audio: false rejected | No audio references |
higgsfield-genjutsu | No generate_audio and no audio references | None documented |
seedance-2.5 reference | Audio generated per model capability; audio reference needs an image or video reference | generate_audio, reference_audio_urls |
What this means for a voice plan
A replaced person in Recast speaks with the source's voice. If the source performer's voice is right for the new character, that is a feature. If it is not, the voice has to change at the source: record the line with the intended voice and swap the face afterwards. Sume's TTS and lip-sync surfaces are separate tools for a still image with audio, with their own windows, and are not part of Recast.
Kling motion control is the one with a switch. Setting keep_original_sound to false returns a silent clip, which suits a case where you will lay your own track. The request below shows it.
curl -X POST https://api.sume.com/v1/kling/3.0/motion-control \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: kling-silent-001" \
-d '{
"image_url": "https://example.com/character.png",
"motion_video_url": "https://example.com/motion.mp4",
"duration_seconds": 8,
"keep_original_sound": false,
"mode": "async"
}'Checking the audio you got
Do not assume. After a job finishes, probe the output with video inspect, which reports whether the clip has an audio track in its probe. A frames: false call returns just probe facts, so it is the cheap way to check. If you plan to caption the clip, remember the captions endpoint transcribes the audible speech, and a silent clip fails as caption_no_speech.
Also check length. A kept audio track that no longer lines up with the picture is a sign the source was cut before the swap. Recast and the face-swap pipeline keep the source's timeline, so a mismatch usually comes from an edit you made upstream rather than from the model.
A practical test before a batch: run one short clip through the tool you picked, probe the output, and listen to it once on the device your audience uses. Audio problems are cheap to catch on one clip and expensive to catch on forty. Record in your project notes which row of the table applied, so the next person does not have to rediscover that Genjutsu has no audio control or that Omni will always return sound.
If you need a specific voice over a swapped performance, treat it as two jobs. First produce the picture with the tool that fits the visual brief. Then replace the track in your editor, or generate speech separately and lay it over the clip. Keeping the stages separate also lets you re-do the voice without paying for another video generation, since the video job and the voice job are billed independently.
Rights and disclosure
The source audio belongs to whoever recorded it, and the swapped face belongs to whoever you photographed. Keep consent for both with the project. Platforms such as TikTok and YouTube have rules about altered or synthetic content; Sume's docs describe the tools, not what a platform requires of you, so read the platform's policy for the channel you publish on.
Finally, remember that sound is part of the deliverable's disclosure story. A clip with a kept original voice and a swapped face is a different thing to a viewer than a clip with a fully synthetic voice, and your own notes should say which one you shipped.
- Plan the voice before the swap.
- Probe the output for an audio track.
- Keep consent for both the performance and the replacement face.
Sources
Related posts
More in Models
- Where to try MAI-Voice-2.1 before paying: a 10-minute listening test
Microsoft lists the MAI Playground, Copilot Audio Expressions and Foundry for MAI-Voice-2.1. Run a ten-minute listening test with a fixed script and cost it.
- Which AI video models accept an input video on Sume?
Seedance, Wan 3.0, H3, H3 Max and Gemini Omni Flash take video references; Recast, Genjutsu and Omni edit need a source video. Kling and Grok take none.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models accept quality? Only five do
Only five Sume image catalog rows list a quality field: GPT Image 2, 2.5 and Sunburst, Ideogram V3 and 4.5. The rest return 400 unsupported_parameter.
Written by Sume