Search footage for a spoken phrase with an API: words with timing
DaVinci Resolve 21 lists IntelliSearch. To find a spoken phrase in a clip by API, transcribe with video inspect and search words[] in your own code.

To find where a phrase is spoken in a clip, run POST /v1/video-inspect with transcribe: true, then search the returned words[] in your own code and use the match time to pull a still. The docs list a transcript and stills, and no search endpoint over footage.
Blackmagic's DaVinci Resolve 21 page lists "Search Content with AI IntelliSearch" among its AI tools. This post covers finding spoken words in one clip with the public routes, from the Video inspect docs, read 2026-09-30.
What does the transcript give me to search?
With transcribe: true the resource carries transcript with text, words[], optional sentence segments[] (set segmentation.mode: "sentence") and an audio_url. Set language_code (for example en or ko) as a hint, or omit it to auto-detect. The rate is $0.01 per audio minute; probe and stills stay unbilled.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: find-phrase-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"language_code": "en"
}'How do I jump from a match to a picture?
The docs name words[] but this page does not list its fields, so read a real response to see how each word's time is named. Take the matched word's time and call POST /v1/video-frames with at: [t] to get a durable still at that instant, at source size. Every at value must satisfy 0 <= t < duration. For a skim, video inspect offers seek: "fast", which snaps each still to the keyframe at or before its instant; keep precise when the timestamp must match.
What are the limits?
The docs cap video inspect at a source of 1800 s, and a silent clip fails with inspect_source_has_no_audio. Check probe.has_audio first with frames: false. This route is about spoken words; the docs show no search over what is on screen.
| Need | Route | Returns |
|---|---|---|
| Spoken words with timing | POST /v1/video-inspect | words[], optional segments[] |
| Still at the match | POST /v1/video-frames | frames[{t,url,width,height}] |
| Is there audio at all | video-inspect with frames false | probe.has_audio |
Sources
Related posts
More in Developers
- Create vs read rate limits: Managed Agents 300/1,200, Sume plans
Anthropic's Managed Agents limit creates and reads separately. Sume splits its per-minute budget the same way by plan, so polling cannot starve submits.
- Submitting 20 avatar videos at once: what each Sume plan accepts
Sume accepts as many paid jobs as concurrency plus queue allows: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Past that a submit returns 429 queue_full.
- Which model does Sume Agent Completions run? Only sume-agent
The model field on POST /v1/agent/completions accepts only sume-agent. Omit it for the same agent; any other value returns 400 invalid_request.
- sume --agent --json: what it redacts before you paste output
Sume CLI --agent mode redacts or summarizes URL-like and account fields where supported. What to still strip before pasting output into a ticket or chat.
Written by Sume