Transcribing a 3-minute Short: duration_seconds 180 reserves $0.03
Video inspect with transcribe true reserves 1 audio minute of STT by default. Send duration_seconds 180 for a 3-minute Short so the hold is 3 x $0.01 = $0.03.

Send duration_seconds 180 when you transcribe a 3-minute Short, so the speech-to-text part of the reservation matches the clip: 3 audio minutes x $0.01 = $0.03. Without the hint, the Video inspect docs say Sume reserves 1 minute ($0.01), and the hint is capped at 600 seconds (as of 2026-10-08). YouTube allows Shorts up to 3 minutes, so 180 is the top of the range.
What the hint does
Inspect with transcribe true runs Sume STT 1.0 on the audio. The public rate is $0.01 per audio minute, and the transcript adds that rate to the Modal compute reservation of the inspect itself. The capture is never more than the hold. The docs do not give a fixed Modal number, so this post prices only the STT part.
| duration_seconds | Audio minutes | STT at $0.01/min |
|---|---|---|
| Not sent | 1 (default reserve) | $0.01 |
| 60 | 1 | $0.01 |
| 180 | 3 | $0.03 |
| 600 | 10 (maximum hint) | $0.10 |
The request
Fields language_code, segmentation and duration_seconds are only valid with transcribe true; sending them alone returns 400 video_inspect_transcribe_required. Adding segmentation.mode sentence returns sentence segments in the shape of caption lines. Add frames false when you want only the probe and transcript without stills.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/short.mp4",
"frames": false,
"transcribe": true,
"duration_seconds": 180,
"segmentation": { "mode": "sentence" }
}Silent clips
A clip with no audio track fails with inspect_source_has_no_audio. Check probe.has_audio first, using an inspect with frames false, as the docs suggest. For a silent Short, skip transcription and write the cues yourself. The source cap for Inspect is 1800 seconds, far above a 3-minute Short.
Why it matters
The default reserve of 1 minute is a hold, not a promise that the transcript is 1 minute long. The docs say the capture is never more than the hold, so an undersized hold protects nothing and an oversized hold ties up credits. Matching the hint to the clip is the cleanest choice: 180 seconds for a full-length Short, 59 for a sub-minute one (still 1 minute, $0.01). If you do not know the duration, run the probe first with frames false and read the duration from it, then send the transcribe call with that figure rounded up. For a batch of 20 full-length Shorts, the STT part of the holds is 20 x $0.03 = $0.60, against 20 x $0.01 = $0.20 with the default. Use the language_code hint when you know the language, as auto-detect is the fallback the docs describe. The sentence segmentation output is shaped like caption lines, which is useful when you feed the text to a caption step later.
Reading the answer
The result carries the probe, with probe.has_audio among its fields, and a transcript object with text, words and, when you asked for sentence segmentation, segments. Read the words array when you need per-word timing for a caption pass, and the segments when you want one caption per line. Neither field changes the price: the rate is per audio minute of the hint, not per word. Keep the transcript with the job id so you do not pay to transcribe the same Short twice.
Sources
Related posts
More in Developers
- Transcribe a 90-second clip with video inspect: cost and reservation
Video inspect with transcribe true adds Sume STT at $0.01 per audio minute; send duration_seconds 90 so the hold fits, or it reserves one minute.
- Transcript from a video on Sume: inspect transcribe or detach first?
Sume can transcribe a video through video inspect (1,800 s limit, compute plus $0.01 per minute) or from a detached 16 kHz mono wav. How to choose.
- TTS word timestamps: timestamps.words and sentence segmentation
Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.
- Turn a roleplay debrief into an avatar feedback clip in Python
Take the written debrief from a roleplay or survey session and render it as a 16:9 Sume avatar clip with a retry-safe key, a 12 to 168 word check and polling.
Written by Sume