Ask Studio says my Short drags: cut the dead air in Python
When feedback says a Short drags, find pauses from word timings and rebuild it without them. Tested Python and the Timeline 1.0 body.

If YouTube's draft feedback says your Short drags, the quickest fix is to remove the silent gaps between spoken words. Get word timings from Sume video inspect, keep the spans where someone is speaking, and rebuild the Short with Timeline 1.0 from those spans. The script below does the arithmetic and prints a render body.
What the feedback gives you
At Made on YouTube 2026, YouTube said Studio feedback gives personalized advice on pacing, structure, storytelling and more (YouTube Blog). Ask Studio is described as a conversational tool for iterating on titles, thumbnails and outlines (YouTube Blog). Neither page gives a threshold for a good pace. Treat the advice as a prompt to look at your own pauses.
Get the word timings
Run POST /v1/video-inspect with frames: false and transcribe: true. The result carries a transcript with words[], the same word timing shape as Sume's STT result, with start and end seconds for each token. Transcription costs $0.01 per audio minute, plus the inspect compute. Detach the audio once with POST /v1/audio-detach ($0.01), because Timeline 1.0 builds its audio from audio.parts[].
The script
keep_ranges merges words whose gap is at most max_gap seconds and pads each span by a tenth of a second. timeline_body turns the spans into matching audio parts and video slots, placing each slot right after the previous one. Both sides use the same source_in, so voice and picture stay together.
import math
def keep_ranges(words, max_gap=0.6, pad=0.1):
spans = [(w["start"], w["end"]) for w in words if w.get("type") != "spacing"]
out = []
for s, e in spans:
if out and s - out[-1][1] <= max_gap:
out[-1][1] = e
else:
out.append([s, e])
return [(max(0.0, s - pad), e + pad) for s, e in out]
def timeline_body(ranges, video_url, audio_url):
parts, video, t = [], [], 0.0
for s, e in ranges:
d = round(e - s, 2)
parts.append({"url": audio_url, "source_in": round(s, 2), "duration": d})
video.append({"source_url": video_url, "start": round(t, 2),
"source_in": round(s, 2), "duration": d})
t += d
return {"audio": {"duration_seconds": math.floor(t), "parts": parts},
"video": video}Check it before you pay
Send the body to POST /v1/timeline-1.0/plan first. It is free and returns duration_seconds, segment_count and billable_minutes. Keep max_gap and pad in mind: a small gap threshold gives a punchy cut but many slots. Timeline 1.0 defaults to render.strategy: auto, which chunks past 12 segments, and it refuses the single strategy above 12 slots. Hard cuts avoid the limit of 8 chained fades.
What the cut can not do
The script cuts only silence. It does not fix structure or storytelling, and it will clip a deliberate pause. Play the result and raise max_gap where the beat matters. If the cut is shorter than a minute, the whole rebuild costs about $0.12 to $0.13: $0.10 for the render, $0.01 for the detach, and about $0.01 for transcription of a minute or less, plus the inspect compute.
| Step | Rate |
|---|---|
| Audio detach | $0.01 |
| Video inspect with transcript | $0.01 per audio minute plus inspect compute |
| Timeline 1.0 render | $0.10 per ceil(output minute) |
Sources
Related posts
More in Media tools
- Audio detach size for a 180 s Short: 34.56 MB wav, 2.88 MB mp3
The byte arithmetic for a 3-minute track from audio detach: stereo 48 kHz wav, mono 16 kHz wav and 128 kbps mp3, with the fields that set each size.
- Can Gemini Omni edit change a video's aspect ratio? No
On Sume, Omni's edit mode takes a video_url and a prompt, defaults to 720p and accepts no aspect_ratio or duration. A 16:9 source stays 16:9, so plan the crop.
- Caption a talking-head video: style, placement and phrasing
Caption a talking-head clip with the Sume API: pick a style, move the line off the face, set words per card. One request, $0.20 for up to 60 seconds.
- Captions for a video with loud background music: script text or cues
Music can bury speech and trip speech-to-text. Sume documents three ways to supply your own wording: script_text, a words array or cues. Which to pick.
Written by Sume