Pull newsletter quotes from a video, with timestamps, via the API
Video inspect with transcribe and sentence segmentation returns timed sentences for $0.01 per audio minute, so a newsletter can link each quote to its moment.

To pull quotable lines from a video with timestamps, call Video inspect with transcribe: true and segmentation: { mode: "sentence" }. You get the transcript plus gapless sentence segments[], shaped like caption lines. Transcription is billed at $0.01 per audio minute; the probe and stills are free.
A newsletter writer can then choose sentences and link each to its moment in the video. Sume does not rank quotes or summarise the talk. It gives you accurate text and time anchors, and you do the choosing.
What do I send?
The clip must already be on media.sume.com in your workspace; import first with POST /v1/media-imports. Idempotency-Key is required. Set frames to false if you only want text, which skips the stills. Omitting duration_seconds reserves one minute, and the hint maxes at 600 seconds, so pass the real length for a longer clip.
mode defaults to sync, waits up to 30 seconds and answers 200 with the finished inspect, or 202 with a queued job to poll. A long talk will probably be 202.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: keynote-quotes-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/keynote.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"duration_seconds": 540,
"segmentation": { "mode": "sentence", "silence_split_seconds": 1 }
}'What does it cost and refuse?
| Item | Value | Why it matters |
|---|---|---|
| Price | $0.01 per audio minute | Confirm live in GET /v1/catalog |
| silence_split_seconds | 0.2 to 3 | Splits a sentence at pauses at least that long |
| Source cap | 1800 seconds | Inspect refuses longer clips |
| Silent clip | inspect_source_has_no_audio | Check probe.has_audio first with frames false |
| STT fields without transcribe | video_inspect_transcribe_required | language_code, segmentation and duration_seconds need transcribe true |
How do I turn segments into linked quotes?
Read the finished resource with GET /v1/video-inspect/:id. The docs describe segments[] as sentence, gapless, caption-line shaped, but I would print one response before wiring code to field names. Then link each chosen sentence to the moment it starts. A ?t= seconds parameter works on most video hosts, but that is a host feature, not something Sume does for you, so check yours.
If you want the quote as a clip rather than a link, Video trim cuts a range from a hosted file for $0.02 and returns a new MP4. Two seconds of lead-in before the sentence start reads better than a hard cut at the first word.
What should I watch for?
Speech-to-text errs on names, product terms and crosstalk. Read every quote against the audio before it goes in an email, especially anything attributed to a person. A language hint (en, ko and so on) helps; omit it to auto-detect.
The transcript lives in the job result. Download what you need, and do not treat the temporary result as your archive. For polling, status and results, see Jobs and results.
Sources
Related posts
More in Use cases
- Nextdoor ad images: 1080x1080 or 2:1, 30 MB, 120-char headline
Nextdoor image ads use 1:1 at 1080 x 1080 or 2:1 at 1080 x 540, JPG or PNG up to 30 MB. Make the 1:1 with Sume and crop a 2:1 locally.
- Ofcom hash matching for AI deepfake intimate images: 30 September 2026
Ofcom said platforms should use hash matching to stop intimate-image abuse, including AI deepfakes, by 30 September 2026. What that means for AI video makers.
- Online course, 20 three-minute avatar lessons: cost by quality tier
Twenty 3-minute avatar lessons are 3,600 seconds. On Sume that is $662.40 Standard, $882.00 Plus or $1,980.00 Max, plus $0.95 once to create the avatar.
- OpenAI TTS requires AI-voice disclosure: burn it with Sume cues
OpenAI's TTS guide says to tell users the voice is AI-generated. A Sume caption job can burn that notice into video with authored cues, no ASR.
Written by Sume