Transcribe a 90-second clip with video inspect: cost and reservation

Video inspect with transcribe true adds Sume STT at $0.01 per audio minute; send duration_seconds 90 so the hold fits, or it reserves one minute.

5 min readSume
All posts

The rate and the reservation

Setting transcribe true on a Sume video inspect runs Sume STT 1.0 on the clip's audio at the public rate of $0.01 per audio minute. For a 90-second clip that is 1.5 minutes at that rate, so about 1.5 cents for the transcript, on top of the inspect's own compute. Send duration_seconds: 90 so the reservation matches; without it Sume reserves one minute. The docs do not state how partial minutes are rounded, so budget two cents to be safe.

Fields that matter

Transcript fields on video inspect (Sume docs, read 2026-10-08)
FieldMeaning
transcribetrue runs STT 1.0 on the clip audio
duration_secondsReservation hint, maximum 600; omit to reserve 1 minute
language_codeHint such as en or ko; omit for auto-detect
segmentation.modesentence returns gapless sentence segments
silence_split_seconds0.2 to 3, with segmentation

Errors to expect

Sending language_code, segmentation or duration_seconds without transcribe true returns 400 video_inspect_transcribe_required. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first; a frames false inspect is enough for that probe, and the probe and stills are free.

Source cap for inspect is 1,800 seconds, but the duration hint maximum is 600, so for longer sources detach and split the audio first.

Why sentence segments help

Sentence segments have no gaps and are in the shape of caption lines, so you can feed them to a caption job as cues if you want to burn authored text. Reference: Video inspect and Video captions.

Reserve and capture

Sume reserves a ceiling at submit and captures the actual amount at the end, never more than the hold. The transcript adds its per-minute rate to the reservation, so a hint that is too small could under-reserve, and a missing hint holds one minute. Sending the true length keeps the hold honest.

Default mode is sync with a 30-second wait: you get the finished inspect if the box answers in time and a queued job otherwise, so handle both cases in code.

Using the sentence segments

Request segmentation.mode sentence and you get segments with no gaps, which is the shape of caption lines. For a 90-second clip you might get a dozen segments. You can feed them as cues to a caption job, but if you only want captions of the speech, a caption job without cues runs its own speech-to-text for $0.20, so compare the two paths before paying twice.

Sources

More in Developers

All Developers posts

Written by Sume