hypit boundaries: cut candidates and the 0.1 threshold
hypit boundaries scores possible cuts from a 12 fps 32x32 difference pass. What rate and threshold do, and how it differs from reference ingest's shot list.

hypit boundaries returns candidates[] {at, score}: possible cut times scored from a 12 fps, 32x32 mean-absolute-difference pass, with defaults rate 12 and threshold 0.1. It is a mechanical series to read alongside frames, not a verdict on where the edits are.
It is unbilled and, like the rest of the lane, flag-gated: it is listed where the matching environment flag allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.
How is this different from reference ingest shots?
Reference ingest's shots[] is the finished answer: frame-exact cuts from ffmpeg scdet and PySceneDetect voting together, tiling [0, duration] with no gap. hypit boundaries is raw material: scored candidates for an agent that is reading the clip in stages, with tiles and cuts to confirm them.
| Reference ingest | hypit boundaries | |
|---|---|---|
| Output | shots[] tiling the clip | candidates[] {at, score} |
| Method | scdet + PySceneDetect vote | 12 fps 32x32 mean abs. difference |
| Clip cap | 300 s | 300 s |
| Billing | Unbilled | Unbilled |
What do rate and threshold change?
rate is how many frames per second are compared (default 12). threshold is the score a difference must reach to appear (default 0.1). Lower the threshold and you see more candidates, including fades and fast motion that are not cuts; raise it and a soft dissolve may vanish. The docs give no tuned values beyond the defaults, so start with them and verify each candidate on a tile.
curl -X POST https://api.sume.com/v1/hypit-understand/boundaries \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hypit-bd-001" \
-d '{"understanding_id":"<probe job id>","rate":12,"threshold":0.1}'When would I pick boundaries over ingest?
When you are already inside a hypit read of the clip and want the cut series next to your transcript and tiles, or when you want to see near-threshold candidates that a voting detector would drop. When you just need a clean shot list to drive a remake, reference ingest gets there in one call and tiles the whole clip with no gap.
Neither tells you why an edit exists. The cut series is timing; the reason a cut lands on a particular word is your reading of the tile and the transcript together.
How do I confirm a candidate?
Run tile with at[] set to the candidate times, or every_frame for a window of 4 seconds or less, which returns every native frame. Then use cut to burn source time into an excerpt (label_time) if you want to review it as video.
There is no stage that depends on the transcript here: boundaries, overview tiles and dense tiles can run in parallel containers right after the probe.
Sources
Related posts
More in Media tools
- Word-level transcript with confidence scores: hypit transcribe
hypit transcribe returns each word with start, end and a 0-1 score using WhisperX, or Sume STT 1.0 without scores. Engines, checkpoints, billing and errors.
- Keyframe trim starts early: fix it with actual_start_seconds
A video-trim with precision keyframe can begin a GOP before your start time. The result reports actual_start_seconds; use it to re-base the next step.
- Korean caption styles compared: weight-shift to editorial-emphasis
Sume has six Hangul caption identities plus korean-ad. Compare black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis.
- Korean subtitle fonts for a video API: 29 faces on Sume
Sume's caption font field names 29 Hangul faces, all SIL OFL 1.1, from Pretendard to Gmarket Sans. Which suit which style, and what fails on a Latin style.
Written by Sume