Video inspect transcribe: no duration hint reserves 1 minute
Sume video inspect with transcribe true bills $0.01 per audio minute. With no duration_seconds it reserves 1 minute; the hint maxes at 600 s (10 min = $0.10).

When you send transcribe: true to Sume video inspect, the transcript is billed at the speech-to-text rate of $0.01 per audio minute. Without a duration_seconds hint Sume reserves 1 minute, and the hint can be at most 600 seconds, so a 10-minute clip reserves 10 x $0.01 = $0.10.
The rule is on the video inspect page, and the live rate is in the catalog. duration_seconds is allowed only together with transcribe: true; otherwise the request returns 400 video_inspect_transcribe_required.
Reserve by hint
The table applies the $0.01 per minute rate to the hint. Whether the final charge matches the reserve is described by the job result, so read it there.
| duration_seconds hint | Minutes | Arithmetic | Reserve |
|---|---|---|---|
| Omitted | 1 | 1 x $0.01 | $0.01 |
| 120 | 2 | 2 x $0.01 | $0.02 |
| 300 | 5 | 5 x $0.01 | $0.05 |
| 600 (maximum) | 10 | 10 x $0.01 | $0.10 |
Clips longer than 10 minutes
Inspect accepts sources up to 1800 seconds, but the hint stops at 600. For a 28-minute clip, cut it with video trim first, in ranges of up to 10 minutes, so each pass matches a hint. Three trims at $0.02 plus three transcripts at up to $0.10 is 3 x $0.02 + 3 x $0.10 = $0.36 at the maximum, which is the figure to budget even if the real charge is lower.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"duration_seconds": 300
}Options that come with transcribe
language_code is an STT hint such as en or ko, and auto-detect is used when you omit it. segmentation.mode: "sentence" splits by sentence, with silence_split_seconds from 0.2 to 3. A clip with no audio track returns inspect_source_has_no_audio, so read probe.has_audio first.
- The default mode is sync with a 30 second wait.
- The source clip must be on your workspace's media host.
language_code,segmentationandduration_secondswithouttranscribe: truereturn a 400.
Why the hint matters
The reserve is a hold placed before the work, so a hint that matches the real length keeps the hold small and the estimate honest. With no hint the hold is one minute, which under-reserves a 10-minute clip; with a hint of 600 the hold is $0.10. Send the real length when you know it, for example from the probe in an earlier call.
If you are running many clips, add the figures. Twenty 5-minute clips are 20 x 5 x $0.01 = $1.00 in transcripts. The same twenty without hints would reserve 20 x $0.01 = $0.20 up front, which is why a hint gives a truer picture of what the job needs.
Choosing sentences or raw text
If you need the transcript as sentences with times, set segmentation.mode to sentence. Silence of silence_split_seconds also ends a sentence, between 0.2 and 3 seconds. A smaller value cuts more often, which suits fast speech and short captions; a larger value keeps long thoughts together. Pick the value by reading a sample before you batch.
Sources
Related posts
More in Media tools
- Video inspect transcript for an 8-minute clip: $0.08 plus compute
transcribe true in video inspect adds STT at $0.01 per audio minute. An 8-minute clip is $0.08 plus compute; duration_seconds hints up to 600 s of reserve.
- Video upscale API: a 12-second clip is $0.108, the cap is 30 seconds
Sume Video Upscale 1.0 bills $0.009 per input second, up to 30 seconds ($0.27). Factor 1.1 to 4, three tiers, and a 5-second default reserve.
- Vidu Q4 540p draft to 1080p: Sume video upscale at 2x costs $0.09
A 10-second 960x540 draft upscaled at the default factor 2 gives 1920x1080 for $0.09 on Sume. Sume does not list Vidu; the clip only needs a public HTTPS URL.
- Voice greeting on a still photo: TTS first, then H3 Max lip sync
To make a still photo speak, make 5 to 14.8 seconds of TTS audio, then call MiniMax H3 Max lip sync. TTS costs $0.0475 per 1,000 characters.
Written by Sume