How many words fit a 5 to 14.8 s lip-sync segment: measure first
At the 2.8 words a second Sume's avatar planner assumes, 5 to 14.8 s is 14 to 41 words. TTS pace varies: measure the audio before setting duration_seconds.

At the 2.8 words per second that Sume's avatar script planner assumes, the lip-sync window of 5 to 14.8 seconds holds about 14 to 41 words. That figure is only a planning guess for writing a script. The H3 Max lip-sync route does not take text; it takes an audio file and a duration_seconds, so the real length is what your text-to-speech voice produces, and you should measure it before you submit.
The estimate
The 2.8 words per second comes from the avatar-workflows package in the Sume repo, which uses it to plan Avatar 1.0 talking videos. It is not part of the lip-sync doc, so treat it as a starting point. Voices differ in pace, and punctuation adds pauses.
| Segment length | Words at 2.8 a second | 768p cost |
|---|---|---|
| 5 s | 14 | $0.50 |
| 8 s | 22 | $0.80 |
| 10 s | 28 | $1.00 |
| 12 s | 33 | $1.20 |
| 14.8 s | 41 | $1.50 (billed as 15 s) |
Measure the audio
Check the length of the real file before every submit. This function reads a WAV file with the Python standard library and rejects a length outside the window.
import wave
def seconds(path):
with wave.open(path, "rb") as w:
return w.getnframes() / w.getframerate()
def duration_for_lip_sync(path):
s = seconds(path)
if not 5 <= s <= 14.8:
raise ValueError("%.2f s is outside 5 to 14.8" % s)
return round(s, 1)
Steps
- Write the line at roughly 14 to 41 words.
- Create the audio and measure it, rather than trusting the estimate.
- Round the length to one decimal and send it as
duration_seconds. - If the number is below 5 or above 14.8, edit the text and render again.
- Submit the job with an
Idempotency-Key.
What changes the pace
Several things move the words-per-second figure. A voice that reads slowly, with long pauses at commas, can run at two words a second, and a brisk voice can run at three and a half. Numbers and abbreviations expand when spoken, so a year such as 2026 is more than one word of audio. Names and product terms add small pauses. A line of 41 words could therefore run to 20 seconds in one voice and 12 in another. This is why the lengths in the table are a writing guide only. If you use several voices in one project, keep a short log of measured seconds per word for each, and use it for the next script. Over a few runs it becomes a better estimate than the planner constant.
A note on rounding
Rounding to one decimal can push a 14.84 second file to 14.8, which the route accepts, while the audio is slightly longer than the number you sent. Avoid the edge: if the file is above 14.7 seconds, trim a word. The route also bills ceil of the number you send, so 14.8 bills as 15 seconds.
What Sume does not do
The route does not read your text and does not tell you how long it will run. It also does not clamp: a length of 15 is rejected. Your own measurement is the control here.
Sources
Related posts
More in Media tools
- Instagram Reels borders and logos: crop them off with video filter
Instagram lists Reels with borders, logos or watermarks among those shown less. Crop a border off with Sume video filter, using fractions, and check it free.
- Is a 180 second video still a YouTube Short? Trim to 179.5 for margin
YouTube's Shorts page says up to 3 minutes. It says nothing on rounding, so cut to 179.5 s with Sume video trim (duration accepts decimals, $0.02 a job).
- Is my clip ready for face swap? Preflight with reference ingest
Face swap wants a 4 to 15 second source with usable audio. Read the clip first with reference ingest purpose face_swap and check duration and audio.silent.
- Join 20 voice lines: audio.parts in the render, or a $0.01 audio job
Timeline 1.0 takes up to 20 audio.parts in one render. A Timeline audio job joins up to 20 parts for $0.01 and returns offsets. Which to use; 45 lines.
Written by Sume