Fit a voiceover to a 30-second slot with Sume TTS speed

Measure the first take, divide by the slot length, and set generation_config.speed (0.6 to 1.5). Why a big speed-up is better solved by cutting the script.

5 min readSume
All posts

To fit narration into a fixed slot, generate once at the default speed, read the duration of the result, divide it by the slot length, and send that ratio as generation_config.speed on a second job. Sume's schema accepts speed from 0.6 to 1.5. A 34-second take for a 30-second slot needs about 1.13. In my experience anything past roughly 1.2 starts to sound rushed (a rule of thumb, not a Sume limit), so trim the script instead of squeezing the audio.

The calculation

Speed is a multiplier on pace, so required speed is measured seconds divided by target seconds. This helper clamps to the schema range and warns when the ratio is high:

def speed_to_fit(measured_seconds, slot_seconds, lo=0.6, hi=1.5):
    needed = measured_seconds / slot_seconds
    if needed > 1.2:
        print("Over 1.2x: shorten the script instead.")
    return min(max(needed, lo), hi)

What the first take tells you

Request timestamps: { words: true } on the first take and read the end of the last word, or measure the downloaded audio file. Translated scripts are the usual cause: the same message often takes a different number of seconds in another language, so a slot that worked in English may not in Spanish or German. Rerun the helper per language rather than reusing one speed.

Other pacing controls across vendors

Pace is controlled differently on each service, which matters if you move between them.

Pace and style controls named on vendor pages (read 2026-10-04)
ServiceControlNote
Sume TTS 1.0generation_config.speed 0.6 to 1.5, plus volume and an emotion guideNumeric; repeatable
OpenAI gpt-4o-mini-ttsinstructions for accent, emotion, intonation, speed, tone, whisperingNatural-language steering
Microsoft MAI-Voice-2.1Granular emotion control on both variantsSpeed control not described on the page read

Cost of getting it right

A retake of 500 characters is about 2.4 cents at $0.0475 per 1,000 characters, so a measure-then-fit loop is cheap. The job's duration cap is also worth remembering: synthesized audio over 1200 seconds fails with tts_duration_exceeded. For fitting a whole video, the Timeline 1.0 program lets each scene's duration follow its audio, which is often easier than speeding the voice up. See the jobs and results page for reading the finished job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume