Pause threshold test: 0.4 s versus 1 s silence splits on one clip

Before First Draft or your own trim removes pauses, test what a pause is: run video inspect with silence_split_seconds at 0.4, 0.7 and 1.0 on one clip.

4 min readSume
All posts

Run the same clip through video inspect three times with silence_split_seconds at 0.4, 0.7 and 1.0, then compare how many segments each gives. Cutting every pause makes speech sound rushed, and cutting none leaves dead air, so the right threshold is a choice you make by looking at your own voice.

TechCrunch reports that First Draft cuts pauses and trims clips for a first pass in under ten seconds (TechCrunch, read 2026-10-07). It does not say how long a pause has to be to get cut, so it is worth knowing what your own number is.

The test

Import one talking clip with POST /v1/media-imports. Then call video inspect three times. silence_split_seconds accepts 0.2 to 3 and only works together with transcribe: true; if you send it without that you get 400 video_inspect_transcribe_required.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pause-test-04" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false,
    "transcribe": true,
    "segmentation": { "mode": "sentence", "silence_split_seconds": 0.4 }
  }'

What to record

Use a new idempotency key for each threshold, or the second call replays the first. Write the results down in a table like this one, filled in with your own numbers.

Pause-threshold worksheet (read 2026-10-07)
silence_split_secondsSegments returnedShortest segmentSounds rushed?
0.4your countyour valueyes or no
0.7your countyour valueyes or no
1.0your countyour valueyes or no

From threshold to cut

Pick the value where segments match how you actually pause between thoughts. Then use the segment start and end times as the start and end of video trim jobs ($0.02 each). Trim defaults to exact precision, a frame-accurate re-encode; keyframe is a stream copy that can start up to a group of pictures early and reports actual_start_seconds.

The speech-to-text part is $0.01 per audio minute, plus the compute for the inspect itself. Three passes on a one-minute clip is therefore a few cents plus compute, which is cheap enough to do once per speaker.

What to prepare

Keep the threshold you choose in one config value. A pause rule that lives in the code is a rule you can change in one place the day you decide the cuts feel too tight.

Reading the three results

A low value gives many short segments, and a high value gives fewer, longer ones. Neither is correct. Look for the value where the segment boundaries land where a listener would take a breath. If a segment ends in the middle of a thought, the threshold is too low for your speaking style; if two sentences sit in one segment, it is too high.

Test two speakers if you have them. A fast speaker and a slow speaker may need different values, and a single global threshold is often wrong for both.

Carrying the number forward

Once you pick a value, write it next to the clip type, such as talking head or voiceover. Reuse it for the next ten Reels before you retest. Changing thresholds every time makes your Reels inconsistent, and consistency is what a viewer notices in a series.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume