AI video with native audio vs a scripted voiceover: exact words
Seedance 2.5, MiniMax H3 and Kling make sound with the picture. When a line must be word-exact, generate the voiceover with Sume TTS and join it on a timeline.

If the spoken line in your video must be word-exact (a price, a legal phrase, a product name), do not leave it to a video model's native audio. Generate the voiceover with TTS from your approved script, then combine it with the clip on a timeline. If the audio only has to fit the mood (ambience, a reaction, a few loose words), native audio is faster and cheaper. The new models make that choice real: the Magic Hour tracker lists Seedance 2.5 as an audio-video model up to 30 seconds, MiniMax H3 with native stereo audio at 15 seconds, and Kling VIDEO 3.0 with native audio at 15 seconds (read 2026-10-06).
Sume's Video Router lists seedance-2.5 (4 to 30 seconds) and minimax-h3-max (5 to 15 seconds, native stereo audio with no toggle). The H3 row has no audio toggle, so you cannot ask for a silent version. Read the capabilities of each model in the catalog before you plan, because limits differ per row.
Three checks
Native audio is a side effect of the video prompt. You can steer it, but you do not control the exact text, the pronunciation of a brand name or the timing of a sentence. Three checks tell you which path to take.
- Would a wrong or reworded sentence cause a refund, a legal problem or a support ticket? Then script it.
- Does the line have to match captions that you already approved? Then script it, and use the script as the caption source with
script_text. - Is the sound only atmosphere, or a short interjection? Then take native audio and skip the extra step.
The scripted path
The scripted path has three calls. First, TTS: POST /v1/tts-router/generate with your transcript, a ready avatar voice and timestamps: { words: true } so you know where each word lands. Second, the video, with a model of your choice. Third, a Timeline render that puts the voiceover on the clip. Timeline audio, a separate call at a flat $0.01 per job, joins up to 20 audio parts in the sample domain, so you can concatenate a voiceover with a music bed without drift (docs).
TTS costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default (pricing, read 2026-10-06). A 25-second voiceover is a few hundred characters, so the voice is the smallest line on the bill. The video seconds are the big one.
| Job | Native audio | Scripted voiceover |
|---|---|---|
| Brand name or price said out loud | Risky: wording is model-chosen | Use: your text, spoken as written |
| Captions must match the voice | Needs speech-to-text first | Use: words come with the TTS job |
| Ambient street or crowd sound | Use: made with the picture | Add a music or sound bed on the timeline |
| Same line in six languages | Re-prompt each clip | Use: one clip, six TTS jobs |
| Fastest first draft | Use: one call | Two or three calls |
The middle path
There is a middle path. If you want a real person's face to speak a scripted line, that is lip sync rather than native audio: the H3 Max lip-sync endpoint takes a Sume-hosted audio file of 5 to 14.8 seconds and is billed per second at 480p, 768p or 1080p, with the 1.25 margin applied. TTS output is the natural input. See replace or remove audio in a video for the replace step and fit a voiceover to a slot when the line is too long for the clip.
None of this needs a claim about which model sounds best. Sume's catalog states what each row accepts; your approved text decides the rest.
A practical habit for a team: keep the approved script as the single source. Feed it to TTS for the voice, to script_text for captions and to your reviewer for sign-off. When legal changes a sentence, you edit one file, rerun three jobs, and every output agrees. With native audio you would regenerate the clip and hope the new take says the new words. That is the real cost difference between the two paths, and it shows up in the second revision, not the first.
Also watch the length. A Seedance 2.5 clip can run up to 30 seconds, a voiceover for it can too, and a single TTS request is capped at 1,200 seconds of audio and 20,000 characters, so the limit is never the voice. The slot is. If the read runs long for the clip, measure the finished audio and adjust speed within 0.6 to 1.5 before you cut the script.
Sources
Related posts
More in Use cases
- One Sume TTS job: 20,000 characters is about 26 minutes of narration
At 750 characters per minute, a 20,000-character job covers about 26.7 minutes for 95 cents. Compare Shorts at 3 min, TikTok at 10 min and Reels at 20 min.
- Outbrain outstream video ad: 60 s, 250 MB, a Sume avatar clip
Outbrain outstream video takes up to 60 seconds and 250 MB in 9:16, 1:1 or 16:9. A Sume avatar job fits the length and the ratios. Check the file before upload.
- Photo to watercolor or pencil sketch with an API, composition kept
Turn a photo into a watercolor or pencil drawing on Sume: send it as an input reference to GPT Image 2.5 and say what to keep. There is no strength slider.
- Pinterest visual search ads: how many image variants per call?
Sume's image API takes n up to 10 per call, with lower per-model ceilings. Read the n range from the endpoint record before you plan a variant set.
Written by Sume