Voiceover longer than the video: speed up, trim or lengthen picture

A Sume TTS voiceover runs 35 seconds over a 30-second video. Compare speed up to 1.5x, a shorter script or a longer Timeline slot, with cost and a Python check.

5 min readSume
All posts

When a Sume voiceover is longer than the video, you have three fixes: speed the read up (generation_config.speed runs from 0.6 to 1.5), cut the script and retake, or lengthen the last picture slot in Timeline so the picture covers the voice. Pick by the gap. A gap of 0.5 seconds or less needs nothing, because Timeline lets picture coverage trail the audio by that much. A gap up to about 1.5x is a speed fix. Beyond that, change the script or the picture.

The speed range and per-character price come from the Sume API reference and API pricing; the slot and coverage rules from Timeline 1.0 and the polling steps from Jobs and results, all read on 2026-10-03. Measure first: request timestamps.words and read duration_seconds from the finished job rather than estimating from characters.

Which fix costs what?

Text to speech is billed per transcript character at $0.0475 per 1,000, and speed is not a billing input, so a faster retake costs the same as the first take. Timeline renders at $0.10 per output minute, rounded up. The table assumes a 900-character script.

Three ways to fit 34.8 seconds of voice into a 30-second video, from Sume's pricing and Timeline docs, read 2026-10-03.
FixWhat changesExtra costRisk
Speed up to about 1.16Retake with a higher speedAbout $0.043 for a 900-character retakeFaster read; docs give no exact speed to seconds ratio, so re-measure
Shorter scriptRetake with fewer charactersLess than a full retakeNeeds a rewrite
Longer pictureExtend the last video[] slot by 4.8 sNone beyond the $0.10 render, which stays one minutePicture holds longer; a still is a static hold

How do I decide in code?

Feed the measured voice length and the picture length to a small planner. It reports the gap, a rough speed that would close it if one fits under the 1.5 cap, the cost of a retake, and the render cost. The speed is an estimate only: Sume's docs state the range, not how seconds scale with it, so submit once and read the new length.

import math

def plan(voice_s, video_s, cap=1.5, chars=900):
    gap = voice_s - video_s
    if gap <= 0.5:
        return {"fix": "none", "note": "picture may trail the voice by up to 0.5 s"}
    need = voice_s / video_s
    out = {"extend_last_slot_by_s": round(gap, 2)}
    out["speed_estimate"] = round(need, 2) if need <= cap else None
    out["retake_usd"] = round(chars * 0.0475 / 1000, 4)
    out["render_usd"] = 0.10 * math.ceil(voice_s / 60)
    return out

print(plan(34.8, 30))
print(plan(52.0, 30))

How do I lengthen the picture instead?

Timeline's audio.duration_seconds sets the output length, and video[].duration is the on-screen time of each slot, with video[0].start at 0. Set audio.duration_seconds to the voice length and give the last slot duration equal to that length minus its start. Stills are static holds, so extending one needs no new generation. Run POST /v1/timeline-1.0/plan first; it is unbilled and returns duration_seconds, billable_minutes and estimated_cost_usd_micros. If the voice is the spine, fix the picture and keep the voice.

When should I change the script?

If the planner prints no speed, the gap is over 1.5x and a fast read will sound rushed. Cut the script first, then retake. Counting is cheap: the 900-character retake above costs about four cents, so trim until the measured length fits.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume