Group TTS sentences into 5 to 14.8 second lip-sync segments in Python
A short Python planner that groups sentence durations into segments the Sume H3 Max lip-sync route accepts (5 to 14.8 s) and prints the 768p cost of each.

To feed a long script to the Sume H3 Max lip-sync route, group the sentences into segments that each run 5 to 14.8 seconds, generate one audio file per segment, and submit one job per file. The planner below does the grouping from the sentence durations you measure on your TTS output, and prints the segment lengths and the 768p cost at $0.10 per billed second. It raises an error rather than guess when a sentence cannot fit, because the route rejects out-of-range durations and never clamps them.
The rules it applies
These are the limits from the repo's lip-sync doc. duration_seconds is required and must be from 5 to 14.8. Billing is ceil(duration_seconds) times the rate. A segment is a group of whole sentences, so it never splits a spoken sentence, which keeps the lips in step with the words.
| Rule | Value | Why |
|---|---|---|
| Minimum segment | 5 s | Route rejects duration_seconds under 5 |
| Maximum segment | 14.8 s | Route rejects duration_seconds over 14.8 |
| Billed seconds | ceil of the segment length | Billing basis in the doc |
| 768p rate | $0.10 per second | Fal list $0.08 times 1.25 |
| Split point | Sentence ends only | Never re-cut audio |
The planner
Run it with the durations of your sentences, in order. No network call is made.
import math
def plan(durs, lo=5.0, hi=14.8):
segs, cur = [], 0.0
for d in durs:
if d > hi:
raise ValueError("one sentence is %.1f s: shorten it" % d)
if cur and cur + d > hi:
segs.append(cur)
cur = d
else:
cur += d
segs.append(cur)
if len(segs) > 1 and segs[-1] < lo and segs[-2] + segs[-1] <= hi:
segs[-2:] = [segs[-2] + segs[-1]]
if any(s < lo for s in segs):
raise ValueError("segment under %.0f s: join lines or use Fabric" % lo)
return segs
segs = plan([4.1, 5.2, 3.9, 6.5, 4.4, 5.0, 3.2])
for s in segs:
print(round(s, 1), round(math.ceil(s) * 0.10, 2))
What it prints, and what to do next
For those seven sentences it prints three segments: 13.2 s ($1.40, billed as 14 seconds), 10.9 s ($1.10, billed as 11) and 8.2 s ($0.90, billed as 9), $3.40 in all. If the last group had come out under 5 seconds, the planner would have tried to merge it into the previous one, and would raise an error only when the sum went above 14.8. A single sentence of 15.2 seconds raises at once: shorten it, split it into two sentences, or send that part to Fabric, which the Sume packet guidance recommends for pieces outside the window.
When it succeeds, generate TTS for each group as its own file. Do not cut a longer file at the planned points; the guidance is never to re-cut audio, because the mouth then follows a different waveform than the one that was timed.
What Sume does not do
Sume does not group your sentences or measure your audio for you. The route also does not report an error that suggests a better length; it only rejects the number. The planner is a convenience in your own code, and it does not replace listening to the result.
Sources
Related posts
More in Media tools
- H3 Max lip sync: 480p drafts then a 1080p final for a 6-second line
A 6-second line costs $0.375 at 480p and $1.20 at 1080p on Sume H3 Max lip sync. Drafting pays only past 1.45 expected attempts. The math, with code.
- Half-banner video: a still stacked over a clip for $0.02
Sume timeline compose stacks one still and one video in a single frame for a flat $0.02 per job. Layout ratio, overlay mode, 300 s ceiling and a ready request.
- Holiday music bed under a voice: soundtrack duck_db 0 to 20
Sume Timeline takes a soundtrack bed with gain_db, loop, fade_out_seconds up to 10 and duck_db 0 to 20 under a voice spine. Python plan call and errors.
- How many words fit a 5 to 14.8 s lip-sync segment: measure first
At the 2.8 words a second Sume's avatar planner assumes, 5 to 14.8 s is 14 to 41 words. TTS pace varies: measure the audio before setting duration_seconds.
Written by Sume