Sentence segments from video inspect: caption lines with no gaps
Sume video inspect with transcribe true can return gapless sentence segments, with silence_split_seconds from 0.2 to 3. Feed them to captions as cues.

Set segmentation.mode: "sentence" on a Sume video inspect call with transcribe: true, and the result includes sentence segments[] with no gaps, in the shape of caption lines. You can also set silence_split_seconds between 0.2 and 3 to split on pauses. Those segments map straight onto the cues field of the video captions API, which burns text you supply without running speech-to-text again.
Transcript options
| Field | Value | Note |
|---|---|---|
transcribe | true | Required for any STT field |
language_code | for example en, ko | STT hint; auto-detect if omitted |
duration_seconds | hint, max 600 | Omitted: Sume reserves 1 minute |
segmentation.mode | sentence | Returns gapless segments[] |
silence_split_seconds | 0.2 to 3 | Pause length used to split |
Price
The public STT rate is $0.01 per audio minute (STT_PUBLIC_PRICING), added to the inspect's compute reservation. A 5-minute clip is 5 x $0.01 = $0.05 of STT line cost, and the maximum 600 s hint is 10 x $0.01 = $0.10. Sume captures each inspect at its own container seconds, times Modal list, times 1.25, plus the platform fee, never more than the hold. For scale, Microsoft states $0.54 per hour of audio for MAI-Transcribe-2-Streaming through the end of the year (read 2026-10-08), which is $0.09 for 10 minutes; the two are different products and Sume does not list the Microsoft model.
Errors to expect
Sending language_code, segmentation or duration_seconds without transcribe: true returns 400 video_inspect_transcribe_required. A silent clip with transcribe: true fails with inspect_source_has_no_audio; probe has_audio first with frames: false.
Feeding captions
Take each segment's text, start and end and send them as cues to POST /v1/video-captions. A standalone caption job costs $0.20 for clips up to 60 seconds under the current estimate. You pay once for the transcript and once for the render, and you control the words in between, which is a good place to fix brand names.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: inspect-sentences-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"transcribe": true,
"language_code": "en",
"segmentation": { "mode": "sentence" },
"duration_seconds": 300
}'Sources
Related posts
More in Media tools
- wav or mp3 for joined audio: priming padding at every edge
Sume timeline audio outputs wav by default and mp3 as an option. Keep wav if you join again or lip-sync: mp3 adds priming padding at every edge.
- Which AI video models make a 20-second clip in one job on Sume?
Seedance 2.5 (4-30 s) and Wan 3.0 (2-30 s) cover 20 seconds in one job on Sume. Omni Flash 1.1 stops at 10 s, MiniMax H3 at 15 s. Full duration table.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
Written by Sume