Transcribing a 3-minute Short: duration_seconds 180 reserves $0.03

Video inspect with transcribe true reserves 1 audio minute of STT by default. Send duration_seconds 180 for a 3-minute Short so the hold is 3 x $0.01 = $0.03.

5 min readSume
All posts

Send duration_seconds 180 when you transcribe a 3-minute Short, so the speech-to-text part of the reservation matches the clip: 3 audio minutes x $0.01 = $0.03. Without the hint, the Video inspect docs say Sume reserves 1 minute ($0.01), and the hint is capped at 600 seconds (as of 2026-10-08). YouTube allows Shorts up to 3 minutes, so 180 is the top of the range.

What the hint does

Inspect with transcribe true runs Sume STT 1.0 on the audio. The public rate is $0.01 per audio minute, and the transcript adds that rate to the Modal compute reservation of the inspect itself. The capture is never more than the hold. The docs do not give a fixed Modal number, so this post prices only the STT part.

STT reservation by hint, read 2026-10-08
duration_secondsAudio minutesSTT at $0.01/min
Not sent1 (default reserve)$0.01
601$0.01
1803$0.03
60010 (maximum hint)$0.10

The request

Fields language_code, segmentation and duration_seconds are only valid with transcribe true; sending them alone returns 400 video_inspect_transcribe_required. Adding segmentation.mode sentence returns sentence segments in the shape of caption lines. Add frames false when you want only the probe and transcript without stills.

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/short.mp4",
  "frames": false,
  "transcribe": true,
  "duration_seconds": 180,
  "segmentation": { "mode": "sentence" }
}

Silent clips

A clip with no audio track fails with inspect_source_has_no_audio. Check probe.has_audio first, using an inspect with frames false, as the docs suggest. For a silent Short, skip transcription and write the cues yourself. The source cap for Inspect is 1800 seconds, far above a 3-minute Short.

Why it matters

The default reserve of 1 minute is a hold, not a promise that the transcript is 1 minute long. The docs say the capture is never more than the hold, so an undersized hold protects nothing and an oversized hold ties up credits. Matching the hint to the clip is the cleanest choice: 180 seconds for a full-length Short, 59 for a sub-minute one (still 1 minute, $0.01). If you do not know the duration, run the probe first with frames false and read the duration from it, then send the transcribe call with that figure rounded up. For a batch of 20 full-length Shorts, the STT part of the holds is 20 x $0.03 = $0.60, against 20 x $0.01 = $0.20 with the default. Use the language_code hint when you know the language, as auto-detect is the fallback the docs describe. The sentence segmentation output is shaped like caption lines, which is useful when you feed the text to a caption step later.

Reading the answer

The result carries the probe, with probe.has_audio among its fields, and a transcript object with text, words and, when you asked for sentence segmentation, segments. Read the words array when you need per-word timing for a caption pass, and the segments when you want one caption per line. Neither field changes the price: the rate is per audio minute of the hint, not per word. Keep the transcript with the job id so you do not pay to transcribe the same Short twice.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume