Fit an AI voiceover to a target length with Sume TTS speed 0.6 to 1.5
To hit a target duration, generate once, measure the audio, then set generation_config.speed to measured over target, clamped to Sume's 0.6 to 1.5 range.

To fit a voiceover to a target length on Sume, generate it once at speed 1.0, measure the duration, and rerun with generation_config.speed set to measured divided by target, clamped to the allowed 0.6 to 1.5 range. If the ratio falls outside that range, change the script length instead, because Sume will not stretch audio beyond 1.5 times or below 0.6 times.
Facts about the field come from the Sume OpenAPI contract. The ratio rule is my own arithmetic, and the docs do not promise exact durations, so verify the result.
What the speed field is
generation_config.speed is a number between 0.6 and 1.5 described as a speed multiplier. The older top-level speed field, an enum of slow, normal and fast, is marked deprecated in favor of it. A multiplier above 1 should shorten the clip and below 1 should lengthen it, so the ratio rule follows. The contract does not say the provider hits the multiplier exactly, which is why the second pass should be checked.
| Measured at 1.0 | Target | Speed to send | In range |
|---|---|---|---|
| 33 s | 30 s | 1.10 | Yes |
| 27 s | 30 s | 0.90 | Yes |
| 50 s | 30 s | 1.67 (clamps to 1.5) | No: cut the script |
| 15 s | 30 s | 0.50 (clamps to 0.6) | No: add words |
A helper that clamps the ratio
This helper has no network calls, so you can run it as written. It returns the speed to send and whether the clamp had to act.
def speed_for(measured_s, target_s, lo=0.6, hi=1.5):
if measured_s <= 0 or target_s <= 0:
raise ValueError("durations must be positive")
raw = measured_s / target_s
speed = min(hi, max(lo, raw))
return round(speed, 2), speed != raw
for measured in (33, 27, 50, 15):
print(measured, speed_for(measured, 30))Measure with timestamps, not a stopwatch
Request timestamps.words: true and the finished job returns words[] with start and end seconds, so the last word's end gives the spoken length without downloading the file. Add a little room for trailing silence if your slot is strict. A job whose synthesized audio goes beyond 1,200 seconds fails with tts_duration_exceeded and captures no credit, so very long scripts should be split into several jobs anyway.
Cost of the second pass
Each pass is a separate job. At $0.0475 per 1,000 characters with a one-cent minimum, a 450 character script costs about 2.2 cents, which ceilings to 3 cents. The two passes cost 6 cents together. For a batch, run the first pass at speed 1.0 for every line, compute all ratios, and only rerun the lines that miss the slot.
Listen to the result before shipping, and if the ratio is near a limit, rewriting the line is often the cleaner fix.
Sources
Related posts
More in Developers
- A for await loop over Sume job status: one async generator, 3 runtimes
Wrap Sume status polling in an async generator and read it with for await. One fetch-only file ran unchanged on Node, Bun and Deno and honors the poll hint.
- Generate a Go client from the Sume OpenAPI with oapi-codegen
The full Sume spec trips oapi-codegen on a [number, null] type, but filtering to the operations you use works. Config, generated names and a short caller.
- GitHub Actions concurrency for a Sume render job: the cancel trap
cancel-in-progress stops your workflow, not the Sume job it submitted. Group by branch, cancel via POST /v1/jobs/{id}/cancel in a final step, and cap the run.
- Go 1.27.2 and 1.26.9 patch net/http: update your Sume webhook receiver
Go 1.27.2 and 1.26.9 (2026-10-08) list security fixes in net/http and crypto/tls. A stdlib receiver for Sume webhooks that verifies the signature, in two files.
Written by Sume