Does H3 Max Recast keep the original audio? What to check on Sume

fal says H3 Max Recast preserves the source audio. Sume rejects generate_audio and audio references on it. Confirm a result has sound with video inspect.

5 min readSume
All posts

Yes per the vendor: fal describes H3 Max Recast as recasting the people in a video using reference photos "while preserving the source motion, camera, cuts, and audio". On Sume you cannot change that audio, because the h3-max-recast row rejects generate_audio and accepts video references but not audio. What Sume's docs do not say is that a finished job has an audio track, so check it before you ship.

Everything about the model comes from fal's H3 Max Recast page (read 2026-10-02). Everything about the Sume row comes from the Video generation and Video Router docs.

What can you control about the sound?

Nothing, by design. The Video generation docs list which video models honor audio and video references: the Seedance 2.x models, Wan 3.0, MiniMax H3 and H3 Max take both, while Gemini Omni Flash 1.1, Genjutsu and h3-max-recast accept video references but not audio. A Recast request with generate_audio: true is rejected as invalid input, and the docs say the row takes no audio references.

This is different from the MiniMax H3 text-to-video rows, which always generate native stereo audio. See which video models always add sound for that side. Recast does not generate a soundtrack; it keeps the one you gave it.

Audio on the H3 family rows on Sume (read 2026-10-02)
RowAudio behaviorAudio inputs accepted
minimax-h3Native stereo audio generatedReference audio clips
minimax-h3-maxNative stereo audio generatedReference audio clips
h3-max-recastSource audio preserved, per falNone; generate_audio rejected

How do I confirm the output has audio?

Run a video inspect on the result. Inspect reads one media.sume.com clip in your workspace, and probe facts are unbilled. The docs name probe.has_audio as the field to check; a frames: false inspect is enough, so you pay for no stills and no transcript.

Do this before you build anything on the assumption. Sume's documentation of the Recast row is silent on the soundtrack, and the claim lives on the vendor's page, which can change.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: inspect-recast-out-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/recast.mp4",
    "frames": false
  }'

What if the new person should speak a new line?

Recast will not do it. The source audio carries the original voice and words. Two routes sit outside Recast. The first is MiniMax H3 Max lip sync on Sume, which takes a still plus audio of 5 to 14.8 seconds, covered in the lip sync endpoint post. The second is a different model row that accepts reference audio, such as MiniMax H3 with a reference audio clip.

Neither swaps a person inside an existing video. They make a new clip. If you need the old footage with a new voice, you are stitching, and the audio spine in Timeline is where a replacement track goes.

What does a bad result look like?

If probe.has_audio is false, treat it as a defect in that output: keep the job id and do not retry in a loop, because each attempt reserves another per-second charge at list times 1.25. If the source itself was silent, there is nothing to preserve, so probe the source first.

Probing both ends costs nothing and gives you a clean rule: a source with audio should give a result with audio. That rule is mine, drawn from the vendor's sentence, not a Sume guarantee.

Sources

Related posts

More in Models

All Models posts

Written by Sume