Fit an AI voiceover to a target length with Sume TTS speed 0.6 to 1.5

To hit a target duration, generate once, measure the audio, then set generation_config.speed to measured over target, clamped to Sume's 0.6 to 1.5 range.

5 min readSume
All posts

To fit a voiceover to a target length on Sume, generate it once at speed 1.0, measure the duration, and rerun with generation_config.speed set to measured divided by target, clamped to the allowed 0.6 to 1.5 range. If the ratio falls outside that range, change the script length instead, because Sume will not stretch audio beyond 1.5 times or below 0.6 times.

Facts about the field come from the Sume OpenAPI contract. The ratio rule is my own arithmetic, and the docs do not promise exact durations, so verify the result.

What the speed field is

generation_config.speed is a number between 0.6 and 1.5 described as a speed multiplier. The older top-level speed field, an enum of slow, normal and fast, is marked deprecated in favor of it. A multiplier above 1 should shorten the clip and below 1 should lengthen it, so the ratio rule follows. The contract does not say the provider hits the multiplier exactly, which is why the second pass should be checked.

Speed math for a 30 second slot (arithmetic, not a vendor claim; field limits per Sume OpenAPI, read 2026-10-10)
Measured at 1.0TargetSpeed to sendIn range
33 s30 s1.10Yes
27 s30 s0.90Yes
50 s30 s1.67 (clamps to 1.5)No: cut the script
15 s30 s0.50 (clamps to 0.6)No: add words

A helper that clamps the ratio

This helper has no network calls, so you can run it as written. It returns the speed to send and whether the clamp had to act.

def speed_for(measured_s, target_s, lo=0.6, hi=1.5):
    if measured_s <= 0 or target_s <= 0:
        raise ValueError("durations must be positive")
    raw = measured_s / target_s
    speed = min(hi, max(lo, raw))
    return round(speed, 2), speed != raw

for measured in (33, 27, 50, 15):
    print(measured, speed_for(measured, 30))

Measure with timestamps, not a stopwatch

Request timestamps.words: true and the finished job returns words[] with start and end seconds, so the last word's end gives the spoken length without downloading the file. Add a little room for trailing silence if your slot is strict. A job whose synthesized audio goes beyond 1,200 seconds fails with tts_duration_exceeded and captures no credit, so very long scripts should be split into several jobs anyway.

Cost of the second pass

Each pass is a separate job. At $0.0475 per 1,000 characters with a one-cent minimum, a 450 character script costs about 2.2 cents, which ceilings to 3 cents. The two passes cost 6 cents together. For a batch, run the first pass at speed 1.0 for every line, compute all ratios, and only rerun the lines that miss the slot.

Listen to the result before shipping, and if the ratio is near a limit, rewriting the line is often the cleaner fix.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume