Fit narration to a 30-second slot: tune TTS speed from measured length
Measure a TTS take with timestamps.words, compute the speed that hits 30 seconds, and know when to cut words instead. Python, within Sume's 0.6 to 1.5 range.

Request word timings with the take, read the end time of the last word, and divide: new speed is the old speed times measured seconds over target seconds. A 34.2 second take at speed 1.0 needs about 1.14 for a 30 second slot. Sume TTS speed runs from 0.6 to 1.5, so a take that needs more than that must be rewritten, not sped up.
The formula assumes length scales with 1 / speed. That is our assumption, so confirm the retake's real length.
How do I measure the take?
Send timestamps: { words: true } with the TTS request. The completed result then includes words[] with monotonic start and end seconds. The last word's end is your speech length, which is a better number than the file duration if the file carries silence at the edges.
What speed do I try next?
The function below clamps to the documented range and flags a take that cannot be fixed by speed alone. It runs as written.
LOW, HIGH = 0.6, 1.5 # generation_config.speed range
def next_speed(speed, measured, target):
wanted = speed * measured / target # assumes length scales with 1/speed
return min(HIGH, max(LOW, round(wanted, 2))), wanted
for speed, measured in [(1.0, 34.2), (1.0, 50.0), (1.0, 20.0)]:
clamped, wanted = next_speed(speed, measured, 30.0)
note = "" if clamped == round(wanted, 2) else " <- clamped; edit the script instead"
print(f"{measured}s at {speed} -> try {clamped} (wanted {wanted:.2f}){note}")
When should I cut words instead?
Speech sped to the top of the range sounds rushed, and the clamp case above shows a 50 second take that would need 1.67. Cut the script to roughly the target length instead. The words-per-minute post in the related list gives a quick way to estimate that before spending a take.
Between those extremes, change one thing per retake. Each completed take is a new generation of the same characters, so budget two or three tries per script and stop when the result is within a half second of the slot.
| Measured at speed 1.0 | Wanted speed | Action |
|---|---|---|
| 34.2 s | 1.14 | Retake at 1.14 |
| 30.4 s | 1.01 | Keep, trim 0.4 s in Timeline |
| 50.0 s | 1.67 | Out of range: cut the script |
| 20.0 s | 0.67 | Retake at 0.67 or add a beat |
Does the slot have to match exactly?
Timeline video slots may trail the audio spine by at most 0.5 seconds, so a take within about half a second of the slot can be fitted in the render. Anything longer than that, change the take.
Sources
Related posts
More in Developers
- Fix one sentence in an AI avatar video without a full re-render
ElevenLabs can regenerate only edited dubbing regions. A Sume avatar video is one job, so the fix is to keep clips short and join them: the cost math.
- FLUX 3 pixel-exact local edits vs the Sume mask_url edit
FLUX 3 Image edits several marked elements in one request. Sume's mask_url edit is documented for ChatGPT Image 2.5 only; here is how to run a local edit.
- FLUX API 402, 403 and 503 errors, and Sume's equivalents
BFL returns 402 for credits, 403 for key permission, 503 for load. Sume returns 402 insufficient_credits and 503 provider_capacity_exceeded. What to do.
- Bulk run completed but SKUs failed: retry only the failures
A Sume bulk-run queue is completed when every item is terminal, not when every item succeeded. Read counts.failed, find the failed SKUs and re-queue only those.
Written by Sume