Pick clip lengths from a voice-over script: one sentence per clip

Turn a voice-over into clips: measure each voiced sentence, round up, and pick the models whose duration window contains it. Windows for six Sume models.

5 min readSume
All posts

To pick clip lengths from a voice-over, split the script into one sentence or idea per clip, voice each one, measure the audio file, round the length up to a whole second and send that as duration to a model whose window contains it. The window is what limits the choice. Wan 3.0 takes 2 to 30 seconds, Seedance 2.5 takes 4 to 30, Gemini Omni Flash 1.1 takes 3 to 10, Kling 3 takes 4 to 15 and MiniMax H3 and H3 Max take 5 to 15.

Measure the voiced file and do not guess from the word count. Speaking pace varies by voice and by language, and a guess that is wrong by one second pushes a clip out of a window or wastes a billed second.

The windows

Pick the model by the shot and then check the window, not the other way round. A window decides only whether a model can make the length you need. Everything else, such as whether you need a first frame, references or sound, is a different question answered by the catalog and by the other pages of this series. These come from the Video Router catalog. The duration on /v1/videos is an integer, so a measured 6.4 seconds becomes 7.

Duration windows for text and image to video on /v1/videos (Sume docs, read 2026-10-07)
Model idWindowSmallest clipLargest clip
wan-3.02 to 30 s2 s30 s
seedance-2.54 to 30 s4 s30 s
gemini-omni-flash-1.13 to 10 s3 s10 s
kling-34 to 15 s4 s15 s
minimax-h35 to 15 s5 s15 s
minimax-h3-max5 to 15 s5 s15 s

Split rules that keep clips usable

A sentence is a good unit because a cut can fall at its end, and the viewer expects a change at that point.. Voice-over drives the edit: the viewer hears the sentence end, and a visual change at that moment feels motivated. Clips that change in the middle of a sentence feel like errors.

Work in a sheet with one row per sentence: the text, the measured seconds, the rounded seconds, the chosen model and the clip's own prompt. When you change a line, you change one row and re-voice one file. This also gives you the total runtime, which is the sum of the rounded seconds, before any clip is generated.

  • If a sentence is under the model's floor, pair it with the next one, or pad the clip with a pause and trim in the edit.
  • If a sentence is over the ceiling, split it at a comma, not in the middle of a phrase.
  • Keep consistency: use the same model for the clips of one scene, so the look matches across the cut.
  • Round up, not down: a clip longer than the line can be trimmed, but a clip shorter than the line cannot be stretched.

Pick the model for each clip

The script below measures a voiced file with the standard library, rounds up and lists the models that accept the length. It reads WAV files, so export your voice as WAV.

import wave, math, sys

WINDOWS = {
    "wan-3.0": (2, 30), "seedance-2.5": (4, 30),
    "gemini-omni-flash-1.1": (3, 10), "kling-3": (4, 15),
    "minimax-h3": (5, 15), "minimax-h3-max": (5, 15),
}

def seconds(path):
    with wave.open(path) as w:
        return math.ceil(w.getnframes() / w.getframerate())

for path in sys.argv[1:]:
    n = seconds(path)
    ok = [m for m, (lo, hi) in WINDOWS.items() if lo <= n <= hi]
    print(path, n, "s ->", ", ".join(ok) or "split the line")

Check the table before you commit

A few details are worth handling in the script. If wave cannot read your file because it is an MP3, convert it or read the length with a tool you already use, since the point is only to get the seconds. If a clip lands exactly on a boundary, such as 10.0 seconds, it fits Omni, but 10.2 rounds up to 11 and does not. That one second changes which models are possible, so look at the rounded value and not the raw one.

The windows can change when the catalog changes, so read them from GET /v1/videos/models in your pipeline instead of hard coding the table. The full limits per model are in how long one video request can be. For talking clips with a voice you already have, the window is different: see writing a lip-sync script in beats.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume