Fit a voiceover to a 30-second slot with Sume TTS speed
Measure the first take, divide by the slot length, and set generation_config.speed (0.6 to 1.5). Why a big speed-up is better solved by cutting the script.

To fit narration into a fixed slot, generate once at the default speed, read the duration of the result, divide it by the slot length, and send that ratio as generation_config.speed on a second job. Sume's schema accepts speed from 0.6 to 1.5. A 34-second take for a 30-second slot needs about 1.13. In my experience anything past roughly 1.2 starts to sound rushed (a rule of thumb, not a Sume limit), so trim the script instead of squeezing the audio.
The calculation
Speed is a multiplier on pace, so required speed is measured seconds divided by target seconds. This helper clamps to the schema range and warns when the ratio is high:
def speed_to_fit(measured_seconds, slot_seconds, lo=0.6, hi=1.5):
needed = measured_seconds / slot_seconds
if needed > 1.2:
print("Over 1.2x: shorten the script instead.")
return min(max(needed, lo), hi)
What the first take tells you
Request timestamps: { words: true } on the first take and read the end of the last word, or measure the downloaded audio file. Translated scripts are the usual cause: the same message often takes a different number of seconds in another language, so a slot that worked in English may not in Spanish or German. Rerun the helper per language rather than reusing one speed.
Other pacing controls across vendors
Pace is controlled differently on each service, which matters if you move between them.
| Service | Control | Note |
|---|---|---|
| Sume TTS 1.0 | generation_config.speed 0.6 to 1.5, plus volume and an emotion guide | Numeric; repeatable |
| OpenAI gpt-4o-mini-tts | instructions for accent, emotion, intonation, speed, tone, whispering | Natural-language steering |
| Microsoft MAI-Voice-2.1 | Granular emotion control on both variants | Speed control not described on the page read |
Cost of getting it right
A retake of 500 characters is about 2.4 cents at $0.0475 per 1,000 characters, so a measure-then-fit loop is cheap. The job's duration cap is also worth remembering: synthesized audio over 1200 seconds fails with tts_duration_exceeded. For fitting a whole video, the Timeline 1.0 program lets each scene's duration follow its audio, which is often easier than speeding the voice up. See the jobs and results page for reading the finished job.
Sources
Related posts
More in Developers
- Fit narration to a fixed slot: measure first, then set TTS speed
A 45-second cap or a 30-second slot decides your script. Render once with word timings, compute the speed ratio, and rewrite only if outside 0.6 to 1.5.
- Flare draft, Sunburst final: a two-pass image edit loop on Sume
Find the edit on GPT Image 2.5 Flare at low quality, then send the same request to Sunburst at high for the keeper. Costs and limits from Sume docs.
- format_not_forkable 409 in the Sume API: what it means and the fix
Sume returns 409 format_not_forkable when the id you called names a built-in capability, not a Format card. How to tell, and which ids to call instead.
- Format webhook 3xx: Sume does not follow redirects, use the final URL
A webhook endpoint that redirects counts as a failed delivery attempt. Register the final HTTPS URL, mind www and trailing-slash redirects, then redeliver.
Written by Sume