Measure your voice's characters per second, then budget video slots
Guessing words per minute breaks a 3-second hook. Voice one Sume take, divide characters by duration_seconds, and size each video slot from your own number.

To fit a line into a 3-second slot, voice one real take and divide its character count by duration_seconds from the job result. That number is characters per second for that voice, language and speed, and it replaces any rule of thumb. Sume returns duration_seconds on every completed TTS job, and one 500-character calibration take costs $0.02375 (API reference).
The script
Use text that looks like your real copy, not a pangram. A mix of short and long sentences averages out. Spaces and punctuation count as characters for billing, so count them the same way here.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, json=body, timeout=60,
headers={**H, "Idempotency-Key": key})
r.raise_for_status()
job = r.json()["request_id"]
while True:
s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
if s.get("terminal"):
break
time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
res.raise_for_status()
return res.json()
sample = open("calibration.txt").read().strip()
res = run("/v1/tts-1.0/generate", {
"transcript": sample,
"avatar_handle": os.environ["VOICE_HANDLE"],
"output_format": {"container": "wav", "encoding": "pcm_s16le",
"sample_rate": 44100},
}, "calibrate-v1")
cps = len(sample) / res["duration_seconds"]
print(f"{cps:.2f} characters per second")
for slot in (3, 5, 8):
print(slot, "s slot fits about", int(slot * cps), "characters")Use it as a budget, not a promise
The figure holds for one voice, one language and one speed. Re-measure when you change any of them. generation_config.speed runs from 0.6 to 1.5, and a change there moves the number, so measure again after each change. Treat the slot sizes as a first draft, then voice the real line and check its duration_seconds against the slot.
When the take is too long
You have three levers, in this order. Cut words, which is free. Then raise generation_config.speed a little and listen for strain. Last, move the next shot's start on the timeline. The render takes video[].start as authoritative, and it requires the first slot to start at 0 (timeline docs).
Why measure per voice
Every new voice or model changes the pace. The new releases make that a routine event: Microsoft's Flash tier claims 55 percent faster inference, which says nothing about speaking rate (Microsoft AI, read 2026-10-04). Speed of generation and speed of speech are different numbers. Keep a small table of characters per second by voice and language beside your preset, and update it when you change either one.
A worked example
Suppose a calibration take of 600 characters returns a duration_seconds of 40. That is 15 characters per second for that voice, so a 3-second hook holds about 45 characters and an 8-second product line about 120. These numbers are an example, not a benchmark: your voice, language and speed will give a different figure, which is the reason to measure. Keep the calibration text in your repo and rerun it for $0.0285 whenever a voice or setting changes.
Sources
Related posts
More in Developers
- Midjourney's Aug 29 edit update: spot silent image-model changes
Midjourney improved its V8.2 edit model two days after opening tests. Keep five fixed edit jobs and compare them monthly so a vendor update never surprises you.
- Grok image model retires Nov 2: migrate safely
A pricing tracker says xAI retires legacy grok-imagine-image-quality on Nov 2, 2026. Migrate by pinning model ids and checking the live catalog.
- MiniMax H3 on Sume: generate_audio false is rejected, audio stays on
The minimax-h3 and minimax-h3-max rows always render native stereo audio. generate_audio false returns 400. How to get a silent clip instead.
- MiniMax H3 reference mix: why 9 + 3 + 3 fails the 12 cap
H3 and H3 Max allow 9 images, 3 videos and 3 audios, but only 12 inputs in total. A Python preflight counter and the cost of images past five.
Written by Sume