Find the video frame where a phrase is spoken: hypit tile around
hypit tile with around {phrase} resolves a spoken phrase on the transcript and returns labeled frames around it. Parameters, limits and the error codes.

To see the frames where a phrase is spoken, call Sume's hypit tile verb with around: {phrase, occurrence, padding} and a transcript_job_id. The phrase is resolved on the word-timed transcript, and you get pages of exact frames around that moment, each cell labeled with its time and the words spoken then.
This surface is flag-gated: it is listed where the matching environment flag allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.
What is a tile?
hypit.tile/1 is a set of pages of exact frames. Each cell is labeled HH:MM:SS.mmm and, with transcript_job_id, the words being spoken at that instant with spans and context. Pages are sized so each cell survives the agent host's 1080 px inline downscale, so a 480 px portrait cell stays 480 px with two per page. Pass cell and columns to trade size for count. Tile is unbilled.
How do I choose which frames?
Pick exactly one style of selection.
| Selector | Use |
|---|---|
start/end with every or frames | Even sampling (default about 1.5 per second, at least 4, at most 9) |
at[] | Explicit instants |
around {phrase, occurrence, padding} | A moment resolved from the transcript |
ranges[] | Several stretches in one call |
layout: frames | One native-size labeled picture per instant |
What does a request look like?
Transcribe first, then pass the transcript job id.
curl -X POST https://api.sume.com/v1/hypit-understand/tile \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hypit-tile-001" \
-d '{"understanding_id":"<probe job id>",
"transcript_job_id":"<transcribe job id>",
"around":{"phrase":"link in bio","occurrence":1,"padding":2}}'When is this better than sampling every second?
Phrase lookup is the shortcut when you know what was said but not when. A hook line, a price mention or a call to action each have a moment you want to see: what is on screen as the words land, whether text appears with them, and where the cut falls relative to the sentence.
Sampling every second answers a different question, what the video looks like overall. For that, tile with every or frames, or reference ingest, which already gives one keyframe per shot. The two compose: ingest for the shot list, then a phrase tile for a close read of the hook.
What goes wrong?
hypit_tile_transcript_required is around without transcript_job_id. hypit_transcript_not_found means that job is not a completed transcribe of this understanding in this workspace. hypit_tile_too_many_cells fires when selectors resolve to more than 96 cells; raise every, lower frames or split ranges[]. If the phrase is misheard by the ASR it will not resolve, so check low-score words on the transcript first.
Sources
Related posts
More in Media tools
- Gemini video understanding: 100 vs 300 tokens a second vs Sume
Gemini reads video at about 100 tokens a second, or 300 at high resolution: 60,000 or 180,000 tokens for 10 minutes. How Sume's stills route compares.
- Hook, demo, CTA: assemble a product ad from 3 clips in one render
Build a 15-second holiday product ad from a hook clip, a demo clip and a call-to-action card with one Sume timeline render. Plan first, render second.
- hypit boundaries: cut candidates and the 0.1 threshold
hypit boundaries scores possible cuts from a 12 fps 32x32 difference pass. What rate and threshold do, and how it differs from reference ingest's shot list.
- Word-level transcript with confidence scores: hypit transcribe
hypit transcribe returns each word with start, end and a 0-1 score using WhisperX, or Sume STT 1.0 without scores. Engines, checkpoints, billing and errors.
Written by Sume