Transcribe a 90-second clip with video inspect: cost and reservation
Video inspect with transcribe true adds Sume STT at $0.01 per audio minute; send duration_seconds 90 so the hold fits, or it reserves one minute.

The rate and the reservation
Setting transcribe true on a Sume video inspect runs Sume STT 1.0 on the clip's audio at the public rate of $0.01 per audio minute. For a 90-second clip that is 1.5 minutes at that rate, so about 1.5 cents for the transcript, on top of the inspect's own compute. Send duration_seconds: 90 so the reservation matches; without it Sume reserves one minute. The docs do not state how partial minutes are rounded, so budget two cents to be safe.
Fields that matter
| Field | Meaning |
|---|---|
| transcribe | true runs STT 1.0 on the clip audio |
| duration_seconds | Reservation hint, maximum 600; omit to reserve 1 minute |
| language_code | Hint such as en or ko; omit for auto-detect |
| segmentation.mode | sentence returns gapless sentence segments |
| silence_split_seconds | 0.2 to 3, with segmentation |
Errors to expect
Sending language_code, segmentation or duration_seconds without transcribe true returns 400 video_inspect_transcribe_required. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first; a frames false inspect is enough for that probe, and the probe and stills are free.
Source cap for inspect is 1,800 seconds, but the duration hint maximum is 600, so for longer sources detach and split the audio first.
Why sentence segments help
Sentence segments have no gaps and are in the shape of caption lines, so you can feed them to a caption job as cues if you want to burn authored text. Reference: Video inspect and Video captions.
Reserve and capture
Sume reserves a ceiling at submit and captures the actual amount at the end, never more than the hold. The transcript adds its per-minute rate to the reservation, so a hint that is too small could under-reserve, and a missing hint holds one minute. Sending the true length keeps the hold honest.
Default mode is sync with a 30-second wait: you get the finished inspect if the box answers in time and a queued job otherwise, so handle both cases in code.
Using the sentence segments
Request segmentation.mode sentence and you get segments with no gaps, which is the shape of caption lines. For a 90-second clip you might get a dozen segments. You can feed them as cues to a caption job, but if you only want captions of the speech, a caption job without cues runs its own speech-to-text for $0.20, so compare the two paths before paying twice.
Sources
More in Developers
- Transcript from a video on Sume: inspect transcribe or detach first?
Sume can transcribe a video through video inspect (1,800 s limit, compute plus $0.01 per minute) or from a detached 16 kHz mono wav. How to choose.
- TTS word timestamps: timestamps.words and sentence segmentation
Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.
- Turn a roleplay debrief into an avatar feedback clip in Python
Take the written debrief from a roleplay or survey session and render it as a 16:9 Sume avatar clip with a retry-safe key, a 12 to 168 word check and polling.
- 12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4
A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.
Written by Sume