Pick clip lengths from a voice-over script: one sentence per clip
Turn a voice-over into clips: measure each voiced sentence, round up, and pick the models whose duration window contains it. Windows for six Sume models.

To pick clip lengths from a voice-over, split the script into one sentence or idea per clip, voice each one, measure the audio file, round the length up to a whole second and send that as duration to a model whose window contains it. The window is what limits the choice. Wan 3.0 takes 2 to 30 seconds, Seedance 2.5 takes 4 to 30, Gemini Omni Flash 1.1 takes 3 to 10, Kling 3 takes 4 to 15 and MiniMax H3 and H3 Max take 5 to 15.
Measure the voiced file and do not guess from the word count. Speaking pace varies by voice and by language, and a guess that is wrong by one second pushes a clip out of a window or wastes a billed second.
The windows
Pick the model by the shot and then check the window, not the other way round. A window decides only whether a model can make the length you need. Everything else, such as whether you need a first frame, references or sound, is a different question answered by the catalog and by the other pages of this series. These come from the Video Router catalog. The duration on /v1/videos is an integer, so a measured 6.4 seconds becomes 7.
| Model id | Window | Smallest clip | Largest clip |
|---|---|---|---|
| wan-3.0 | 2 to 30 s | 2 s | 30 s |
| seedance-2.5 | 4 to 30 s | 4 s | 30 s |
| gemini-omni-flash-1.1 | 3 to 10 s | 3 s | 10 s |
| kling-3 | 4 to 15 s | 4 s | 15 s |
| minimax-h3 | 5 to 15 s | 5 s | 15 s |
| minimax-h3-max | 5 to 15 s | 5 s | 15 s |
Split rules that keep clips usable
A sentence is a good unit because a cut can fall at its end, and the viewer expects a change at that point.. Voice-over drives the edit: the viewer hears the sentence end, and a visual change at that moment feels motivated. Clips that change in the middle of a sentence feel like errors.
Work in a sheet with one row per sentence: the text, the measured seconds, the rounded seconds, the chosen model and the clip's own prompt. When you change a line, you change one row and re-voice one file. This also gives you the total runtime, which is the sum of the rounded seconds, before any clip is generated.
- If a sentence is under the model's floor, pair it with the next one, or pad the clip with a pause and trim in the edit.
- If a sentence is over the ceiling, split it at a comma, not in the middle of a phrase.
- Keep consistency: use the same model for the clips of one scene, so the look matches across the cut.
- Round up, not down: a clip longer than the line can be trimmed, but a clip shorter than the line cannot be stretched.
Pick the model for each clip
The script below measures a voiced file with the standard library, rounds up and lists the models that accept the length. It reads WAV files, so export your voice as WAV.
import wave, math, sys
WINDOWS = {
"wan-3.0": (2, 30), "seedance-2.5": (4, 30),
"gemini-omni-flash-1.1": (3, 10), "kling-3": (4, 15),
"minimax-h3": (5, 15), "minimax-h3-max": (5, 15),
}
def seconds(path):
with wave.open(path) as w:
return math.ceil(w.getnframes() / w.getframerate())
for path in sys.argv[1:]:
n = seconds(path)
ok = [m for m, (lo, hi) in WINDOWS.items() if lo <= n <= hi]
print(path, n, "s ->", ", ".join(ok) or "split the line")Check the table before you commit
A few details are worth handling in the script. If wave cannot read your file because it is an MP3, convert it or read the length with a tool you already use, since the point is only to get the seconds. If a clip lands exactly on a boundary, such as 10.0 seconds, it fits Omni, but 10.2 rounds up to 11 and does not. That one second changes which models are possible, so look at the rounded value and not the raw one.
The windows can change when the catalog changes, so read them from GET /v1/videos/models in your pipeline instead of hard coding the table. The full limits per model are in how long one video request can be. For talking clips with a voice you already have, the window is different: see writing a lip-sync script in beats.
Sources
Related posts
More in Use cases
- How do I turn one podcast episode into five quote clips for social?
Transcribe the episode in 10-minute chunks, split five quotes out, put the cover still under each and burn captions: $1.80 for a 28-minute episode on Sume.
- Post-call recap video with an AI avatar for prospects: script and cost
After a sales call, send a 30-second recap clip from a Sume avatar. Script structure, cost at standard, plus and max, and how to keep it honest and reviewed.
- Pre-rendered avatar greetings per visitor segment, not a live avatar
Instead of a live avatar for each visitor, render one short Sume avatar clip per segment ahead of time. Per-tier cost for six 12-second greetings.
- Product photo to a 6-second vertical ad: a still is a static hold
A still in a Timeline video slot is held, not animated: motion is accepted but ignored with a motion_ignored warning. Default 1080x1920, $0.10 for 6 seconds.
Written by Sume