What happens to the audio in each Sume video-to-video tool

Recast and Avatar Face Swap keep the source audio, Kling motion control keeps it by default, Omni always produces native audio, Genjutsu takes no audio field.

6 min readSume
All posts

The audio rule differs by tool, and getting it wrong means a clip that sounds like the wrong person. h3-max-recast keeps the source audio, the Beta Avatar Face Swap muxes the original source audio back, Kling 3.0 motion control keeps the driving video's sound unless you set keep_original_sound to false, gemini-omni-flash-1.1 always generates native synced audio, and higgsfield-genjutsu has no audio field at all (Video Router docs, Videos docs and Face swap docs, read 2026-10-03). Read the table below before you plan a voice.

In every tool here the new person's photo contributes a face, not a voice.

The audio behavior table

Each row is stated in the Sume docs; where a field is absent, the row says so rather than guessing.

Audio in Sume video-to-video and related tools, read 2026-10-03
ToolAudio behaviorControl you have
h3-max-recastSource audio kept in the outputNone; plan the voice upstream
Avatar Face Swap (Beta)Source needs usable audio; original audio is muxed backNone; prompts and provider fields unsupported
Kling 3.0 motion controlDriving video's sound kept by defaultkeep_original_sound: true by default, false gives a silent clip
gemini-omni-flash-1.1Native synced audio always on; generate_audio: false rejectedNo audio references
higgsfield-genjutsuNo generate_audio and no audio referencesNone documented
seedance-2.5 referenceAudio generated per model capability; audio reference needs an image or video referencegenerate_audio, reference_audio_urls

What this means for a voice plan

A replaced person in Recast speaks with the source's voice. If the source performer's voice is right for the new character, that is a feature. If it is not, the voice has to change at the source: record the line with the intended voice and swap the face afterwards. Sume's TTS and lip-sync surfaces are separate tools for a still image with audio, with their own windows, and are not part of Recast.

Kling motion control is the one with a switch. Setting keep_original_sound to false returns a silent clip, which suits a case where you will lay your own track. The request below shows it.

curl -X POST https://api.sume.com/v1/kling/3.0/motion-control \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: kling-silent-001" \
  -d '{
    "image_url": "https://example.com/character.png",
    "motion_video_url": "https://example.com/motion.mp4",
    "duration_seconds": 8,
    "keep_original_sound": false,
    "mode": "async"
  }'

Checking the audio you got

Do not assume. After a job finishes, probe the output with video inspect, which reports whether the clip has an audio track in its probe. A frames: false call returns just probe facts, so it is the cheap way to check. If you plan to caption the clip, remember the captions endpoint transcribes the audible speech, and a silent clip fails as caption_no_speech.

Also check length. A kept audio track that no longer lines up with the picture is a sign the source was cut before the swap. Recast and the face-swap pipeline keep the source's timeline, so a mismatch usually comes from an edit you made upstream rather than from the model.

A practical test before a batch: run one short clip through the tool you picked, probe the output, and listen to it once on the device your audience uses. Audio problems are cheap to catch on one clip and expensive to catch on forty. Record in your project notes which row of the table applied, so the next person does not have to rediscover that Genjutsu has no audio control or that Omni will always return sound.

If you need a specific voice over a swapped performance, treat it as two jobs. First produce the picture with the tool that fits the visual brief. Then replace the track in your editor, or generate speech separately and lay it over the clip. Keeping the stages separate also lets you re-do the voice without paying for another video generation, since the video job and the voice job are billed independently.

Rights and disclosure

The source audio belongs to whoever recorded it, and the swapped face belongs to whoever you photographed. Keep consent for both with the project. Platforms such as TikTok and YouTube have rules about altered or synthetic content; Sume's docs describe the tools, not what a platform requires of you, so read the platform's policy for the channel you publish on.

Finally, remember that sound is part of the deliverable's disclosure story. A clip with a kept original voice and a swapped face is a different thing to a viewer than a clip with a fully synthetic voice, and your own notes should say which one you shipped.

  • Plan the voice before the swap.
  • Probe the output for an audio track.
  • Keep consent for both the performance and the replacement face.

Sources

Related posts

More in Models

All Models posts

Written by Sume