Reel voiceover too long for 60 seconds? Speed setting and re-render
A 68-second TTS read against a 60-second Reel: Sume accepts generation_config.speed from 0.6 to 1.5. The arithmetic, and when a rewrite is better.

If your voiceover runs 68 seconds and the Reel should be 60, ask for a read about 13 percent faster: generation_config.speed on Sume's TTS 1.0 accepts 0.6 to 1.5. The multiplier is the current length divided by the target length, and a regeneration at $0.0475 per 1,000 characters is cheap enough to try twice.
Instagram's Reels guide (read 2026-10-03) calls 15 to 60 seconds the sweet spot for tutorials, day-in-my-life clips and most creative formats, which is where this fit problem usually arises.
The arithmetic
Speed is a multiplier on the pace, so new duration is about old duration divided by speed. This is a planning estimate, because Sume's docs do not publish a guaranteed relation between speed and length; read the real duration_seconds from the new job.
| Current read | Target | Speed to try | Within 0.6 to 1.5? |
|---|---|---|---|
| 62 s | 60 s | 1.04 | Yes |
| 68 s | 60 s | 1.13 | Yes |
| 75 s | 60 s | 1.25 | Yes |
| 90 s | 60 s | 1.50 | At the limit |
| 100 s | 60 s | 1.67 | No: cut words |
Compute it and re-submit
The helper below returns the speed to try, or None when the gap is beyond the accepted range, which is the signal to cut the script instead.
def speed_for(current_s, target_s, lo=0.6, hi=1.5):
s = round(current_s / target_s, 2)
return s if lo <= s <= hi else None
for cur in (62, 68, 75, 90, 100):
print(cur, speed_for(cur, 60))When to rewrite instead
Above roughly 1.25 a read starts to sound hurried to most listeners, so past that point trimming a sentence usually serves the Reel better; that is a listening judgment, not a Sume limit. A script change costs one new TTS job. If you have already cut the voice into parts, regenerate only the part that overran and join again with timeline audio concat, a flat $0.01.
What Sume does not do
Sume does not time-stretch a finished file to fit, and it does not shorten a script for you. Speed is a generation setting: change it and you get a new job and a new duration_seconds. Feed that number into Timeline 1.0 as audio.duration_seconds so the render, and the $0.10 per output minute, match the voice you have.
Sources
Related posts
More in Media tools
- Clipdrop remove-background API: 60 requests a minute vs Sume RMBG
Clipdrop's remove-background API takes a 30 MB, 25 MP upload at 60 requests a minute per key. Sume RMBG takes a public HTTPS image_url and runs as a job.
- Resolve 21.1 multicam up to 25 angles vs Sume timeline slots
Resolve 21.1 adds multicam viewing with shortcuts for up to 25 angles. Sume has no multicam; Timeline 1.0 stacks 1 to 200 clips on one audio spine.
- Resolve 21 IntelliSearch vs a transcript with word times
Resolve 21 IntelliSearch finds moments in your footage. Sume's video-inspect returns a transcript with word times, which you can search for a trim point.
- Resolve 21 Speech Generation vs a text-to-speech API call
Resolve 21 generates speech from text with Blackmagic voice models. Sume's tts_create does it from a script, and a voice-language mismatch returns 409 first.
Written by Sume