AI video with native audio vs a scripted voiceover: exact words

Seedance 2.5, MiniMax H3 and Kling make sound with the picture. When a line must be word-exact, generate the voiceover with Sume TTS and join it on a timeline.

5 min readSume
All posts

If the spoken line in your video must be word-exact (a price, a legal phrase, a product name), do not leave it to a video model's native audio. Generate the voiceover with TTS from your approved script, then combine it with the clip on a timeline. If the audio only has to fit the mood (ambience, a reaction, a few loose words), native audio is faster and cheaper. The new models make that choice real: the Magic Hour tracker lists Seedance 2.5 as an audio-video model up to 30 seconds, MiniMax H3 with native stereo audio at 15 seconds, and Kling VIDEO 3.0 with native audio at 15 seconds (read 2026-10-06).

Sume's Video Router lists seedance-2.5 (4 to 30 seconds) and minimax-h3-max (5 to 15 seconds, native stereo audio with no toggle). The H3 row has no audio toggle, so you cannot ask for a silent version. Read the capabilities of each model in the catalog before you plan, because limits differ per row.

Three checks

Native audio is a side effect of the video prompt. You can steer it, but you do not control the exact text, the pronunciation of a brand name or the timing of a sentence. Three checks tell you which path to take.

  • Would a wrong or reworded sentence cause a refund, a legal problem or a support ticket? Then script it.
  • Does the line have to match captions that you already approved? Then script it, and use the script as the caption source with script_text.
  • Is the sound only atmosphere, or a short interjection? Then take native audio and skip the extra step.

The scripted path

The scripted path has three calls. First, TTS: POST /v1/tts-router/generate with your transcript, a ready avatar voice and timestamps: { words: true } so you know where each word lands. Second, the video, with a model of your choice. Third, a Timeline render that puts the voiceover on the clip. Timeline audio, a separate call at a flat $0.01 per job, joins up to 20 audio parts in the sample domain, so you can concatenate a voiceover with a music bed without drift (docs).

TTS costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default (pricing, read 2026-10-06). A 25-second voiceover is a few hundred characters, so the voice is the smallest line on the bill. The video seconds are the big one.

Which audio path for which job (read 2026-10-06)
JobNative audioScripted voiceover
Brand name or price said out loudRisky: wording is model-chosenUse: your text, spoken as written
Captions must match the voiceNeeds speech-to-text firstUse: words come with the TTS job
Ambient street or crowd soundUse: made with the pictureAdd a music or sound bed on the timeline
Same line in six languagesRe-prompt each clipUse: one clip, six TTS jobs
Fastest first draftUse: one callTwo or three calls

The middle path

There is a middle path. If you want a real person's face to speak a scripted line, that is lip sync rather than native audio: the H3 Max lip-sync endpoint takes a Sume-hosted audio file of 5 to 14.8 seconds and is billed per second at 480p, 768p or 1080p, with the 1.25 margin applied. TTS output is the natural input. See replace or remove audio in a video for the replace step and fit a voiceover to a slot when the line is too long for the clip.

None of this needs a claim about which model sounds best. Sume's catalog states what each row accepts; your approved text decides the rest.

A practical habit for a team: keep the approved script as the single source. Feed it to TTS for the voice, to script_text for captions and to your reviewer for sign-off. When legal changes a sentence, you edit one file, rerun three jobs, and every output agrees. With native audio you would regenerate the clip and hope the new take says the new words. That is the real cost difference between the two paths, and it shows up in the second revision, not the first.

Also watch the length. A Seedance 2.5 clip can run up to 30 seconds, a voiceover for it can too, and a single TTS request is capped at 1,200 seconds of audio and 20,000 characters, so the limit is never the voice. The slot is. If the read runs long for the clip, measure the finished audio and adjust speed within 0.6 to 1.5 before you cut the script.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume