Group TTS sentences into 5 to 14.8 second lip-sync segments in Python

A short Python planner that groups sentence durations into segments the Sume H3 Max lip-sync route accepts (5 to 14.8 s) and prints the 768p cost of each.

4 min readSume
All posts

To feed a long script to the Sume H3 Max lip-sync route, group the sentences into segments that each run 5 to 14.8 seconds, generate one audio file per segment, and submit one job per file. The planner below does the grouping from the sentence durations you measure on your TTS output, and prints the segment lengths and the 768p cost at $0.10 per billed second. It raises an error rather than guess when a sentence cannot fit, because the route rejects out-of-range durations and never clamps them.

The rules it applies

These are the limits from the repo's lip-sync doc. duration_seconds is required and must be from 5 to 14.8. Billing is ceil(duration_seconds) times the rate. A segment is a group of whole sentences, so it never splits a spoken sentence, which keeps the lips in step with the words.

Planner rules and where they come from (read 2026-10-08)
RuleValueWhy
Minimum segment5 sRoute rejects duration_seconds under 5
Maximum segment14.8 sRoute rejects duration_seconds over 14.8
Billed secondsceil of the segment lengthBilling basis in the doc
768p rate$0.10 per secondFal list $0.08 times 1.25
Split pointSentence ends onlyNever re-cut audio

The planner

Run it with the durations of your sentences, in order. No network call is made.

import math

def plan(durs, lo=5.0, hi=14.8):
    segs, cur = [], 0.0
    for d in durs:
        if d > hi:
            raise ValueError("one sentence is %.1f s: shorten it" % d)
        if cur and cur + d > hi:
            segs.append(cur)
            cur = d
        else:
            cur += d
    segs.append(cur)
    if len(segs) > 1 and segs[-1] < lo and segs[-2] + segs[-1] <= hi:
        segs[-2:] = [segs[-2] + segs[-1]]
    if any(s < lo for s in segs):
        raise ValueError("segment under %.0f s: join lines or use Fabric" % lo)
    return segs

segs = plan([4.1, 5.2, 3.9, 6.5, 4.4, 5.0, 3.2])
for s in segs:
    print(round(s, 1), round(math.ceil(s) * 0.10, 2))

What it prints, and what to do next

For those seven sentences it prints three segments: 13.2 s ($1.40, billed as 14 seconds), 10.9 s ($1.10, billed as 11) and 8.2 s ($0.90, billed as 9), $3.40 in all. If the last group had come out under 5 seconds, the planner would have tried to merge it into the previous one, and would raise an error only when the sum went above 14.8. A single sentence of 15.2 seconds raises at once: shorten it, split it into two sentences, or send that part to Fabric, which the Sume packet guidance recommends for pieces outside the window.

When it succeeds, generate TTS for each group as its own file. Do not cut a longer file at the planned points; the guidance is never to re-cut audio, because the mouth then follows a different waveform than the one that was timed.

What Sume does not do

Sume does not group your sentences or measure your audio for you. The route also does not report an error that suggests a better length; it only rejects the number. The planner is a convenience in your own code, and it does not replace listening to the result.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume