TTS then H3 Max lip sync: check the 5 to 14.8 s audio window

H3 Max lip sync takes 5 to 14.8 seconds of Sume-hosted audio and clips the rest. Measure a TTS line from its word timings and price it before you submit.

5 min readSume
All posts

Before you send TTS audio to Sume's H3 Max lip-sync endpoint, measure it. The endpoint accepts a Sume-hosted audio file of 5 to 14.8 seconds, at most 10 MB, and the provider silently clips anything after 14.8 seconds, so a long line loses its ending without an error. You can get the length for free from the same TTS job: ask for timestamps: { words: true } and read the end time of the last word. The script below does that, then tells you whether the line is too short, too long or ready, and what it will cost.

MiniMax H3 is one of the new native-audio video models (15 seconds, stereo audio, open weights since 2026-08-03 per the Magic Hour tracker, read 2026-10-06). Lip sync is the other way to use it: you bring the audio, and the still speaks it. That is the right path when the words must be yours.

Measure the line first

import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]

def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    out = json.load(u.urlopen(u.Request(url, data=data, headers=h)))
    return out.get("data", out)

sub = job = call("https://api.sume.com/v1/tts-router/generate", {
    "model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
    "transcript": "Our spring sale starts Friday, with free shipping on every order.",
    "timestamps": {"words": True}, "output_format": {"container": "wav", "sample_rate": 48000}},
    "h3-line-v1")
while not job.get("terminal"):
    time.sleep(job.get("next_poll_after_seconds") or 2)
    job = call(sub["status_url"])
res = call(sub["result_url"])
res = res.get("result", res)
seconds = max(w["end"] for w in res["words"])
audio = next(a["url"] for a in res["artifacts"] if a["type"] == "audio")
billed = -(-seconds // 1)
if seconds < 5:
    print(f"{seconds:.1f}s is under 5 s: add a sentence or slow the speed")
elif seconds > 14.8:
    print(f"{seconds:.1f}s is over 14.8 s: split the script, the provider would clip it")
else:
    print(audio, f"{seconds:.1f}s, billed {billed:.0f} s, 768p = ${billed * 0.10:.2f}")

Why measure, and how to fix a miss

The guard exists because Sume refuses a duration_seconds outside 5 to 14.8 with invalid_request and never clamps it. A clamped reservation would produce a clip shorter than your audio and lose speech, so the API leaves the decision to you. Short audio is rejected by the provider; long audio is cut at 14.8 seconds. The audio also has to live on Sume's media host, which a TTS artifact does, and fit in 10 MB.

To fix a line that is too long, you have two moves. Raise generation_config.speed toward its 1.5 cap and measure again. Or cut the script at a sentence break and make two clips. If you cut, use segmentation: { mode: "sentence" } on the same TTS request to see where the sentences fall; it needs word timestamps, and gives gapless segments with a default 70 ms boundary lead. For a line that is too short, a slower generation_config.speed down to 0.6 is the cheap move.

What the clip costs

Billing is ceil(duration_seconds) times the per-second list rate for the resolution, times Sume's 1.25 margin: $0.05, $0.08 and $0.16 list at 480p, 768p and 1080p become $0.0625, $0.10 and $0.20 per second, per Sume's H3 Max lip-sync reference. The reservation is captured on completion and refunded on failure. The line in the script is about 60 characters, so TTS adds a fraction of a cent to the $0.10-per-second clip.

H3 Max lip sync price by audio length, 1.25 margin included (read 2026-10-06)
Audio lengthBilled seconds480p768p1080p
5.0 s5$0.32$0.50$1.00
9.4 s10$0.63$1.00$2.00
14.8 s15$0.94$1.50$3.00
15.2 sRejected over 14.8 sn/an/an/a

Where to go next

Sume prices are listed on the API pricing page, and the billing rule above is why 5.1 seconds of audio costs six seconds. Read the rounding explanation in the 5.1-second post if you batch many clips. The job lifecycle for the lip-sync call, after the guard passes, is the same envelope as TTS: status URL, result URL and Idempotency-Key, per Jobs and results.

Two practical notes. The measured length comes from the last word's end time, so it ignores any trailing silence in the file; if your TTS line ends with a long pause, the file can be a little longer than the number you printed, and a line that measures 14.7 seconds is safer trimmed than left at the edge. And because billing rounds up to a whole second, a line of 9.1 seconds bills as 10 seconds, so trimming a breath at the end of a line only saves money when it crosses a whole-second boundary; at 768p each second is $0.10. Keep the measured seconds and the billed seconds as two columns in your log, because the billed second is the one that appears on the invoice.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume