Ask Studio says my Short drags: cut the dead air in Python

When feedback says a Short drags, find pauses from word timings and rebuild it without them. Tested Python and the Timeline 1.0 body.

6 min readSume
All posts

If YouTube's draft feedback says your Short drags, the quickest fix is to remove the silent gaps between spoken words. Get word timings from Sume video inspect, keep the spans where someone is speaking, and rebuild the Short with Timeline 1.0 from those spans. The script below does the arithmetic and prints a render body.

What the feedback gives you

At Made on YouTube 2026, YouTube said Studio feedback gives personalized advice on pacing, structure, storytelling and more (YouTube Blog). Ask Studio is described as a conversational tool for iterating on titles, thumbnails and outlines (YouTube Blog). Neither page gives a threshold for a good pace. Treat the advice as a prompt to look at your own pauses.

Get the word timings

Run POST /v1/video-inspect with frames: false and transcribe: true. The result carries a transcript with words[], the same word timing shape as Sume's STT result, with start and end seconds for each token. Transcription costs $0.01 per audio minute, plus the inspect compute. Detach the audio once with POST /v1/audio-detach ($0.01), because Timeline 1.0 builds its audio from audio.parts[].

The script

keep_ranges merges words whose gap is at most max_gap seconds and pads each span by a tenth of a second. timeline_body turns the spans into matching audio parts and video slots, placing each slot right after the previous one. Both sides use the same source_in, so voice and picture stay together.

import math

def keep_ranges(words, max_gap=0.6, pad=0.1):
    spans = [(w["start"], w["end"]) for w in words if w.get("type") != "spacing"]
    out = []
    for s, e in spans:
        if out and s - out[-1][1] <= max_gap:
            out[-1][1] = e
        else:
            out.append([s, e])
    return [(max(0.0, s - pad), e + pad) for s, e in out]

def timeline_body(ranges, video_url, audio_url):
    parts, video, t = [], [], 0.0
    for s, e in ranges:
        d = round(e - s, 2)
        parts.append({"url": audio_url, "source_in": round(s, 2), "duration": d})
        video.append({"source_url": video_url, "start": round(t, 2),
                      "source_in": round(s, 2), "duration": d})
        t += d
    return {"audio": {"duration_seconds": math.floor(t), "parts": parts},
            "video": video}

Check it before you pay

Send the body to POST /v1/timeline-1.0/plan first. It is free and returns duration_seconds, segment_count and billable_minutes. Keep max_gap and pad in mind: a small gap threshold gives a punchy cut but many slots. Timeline 1.0 defaults to render.strategy: auto, which chunks past 12 segments, and it refuses the single strategy above 12 slots. Hard cuts avoid the limit of 8 chained fades.

What the cut can not do

The script cuts only silence. It does not fix structure or storytelling, and it will clip a deliberate pause. Play the result and raise max_gap where the beat matters. If the cut is shorter than a minute, the whole rebuild costs about $0.12 to $0.13: $0.10 for the render, $0.01 for the detach, and about $0.01 for transcription of a minute or less, plus the inspect compute.

Dead-air rebuild, rates for a Short of up to one minute (read 2026-10-07)
StepRate
Audio detach$0.01
Video inspect with transcript$0.01 per audio minute plus inspect compute
Timeline 1.0 render$0.10 per ceil(output minute)

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume