Video inspect transcript for an 8-minute clip: $0.08 plus compute
transcribe true in video inspect adds STT at $0.01 per audio minute. An 8-minute clip is $0.08 plus compute; duration_seconds hints up to 600 s of reserve.

Setting transcribe: true on video inspect adds Sume STT at $0.01 per audio minute. An 8-minute clip is 8 x $0.01 = $0.08, on top of the probe and stills, which Sume bills by Modal compute. Send duration_seconds: 480 so the reserve matches the clip.
What the hold looks like
Without duration_seconds Sume reserves 1 minute for the transcript. The hint maximum is 600 seconds. The capture never exceeds the hold, per the docs read 2026-10-09.
| Clip | Hint | Arithmetic | STT part |
|---|---|---|---|
| 2 min | 120 | 2 x 0.01 | $0.02 |
| 8 min | 480 | 8 x 0.01 | $0.08 |
| 10 min | 600 (maximum hint) | 10 x 0.01 | $0.10 |
| 25 min | above 600 | 25 x 0.01 | $0.25 (split the audio; see below) |
Request
language_code is an STT hint such as en or ko. segmentation.mode: "sentence" also returns sentence segments without gaps, with an optional silence_split_seconds from 0.2 to 3. These fields without transcribe: true return 400 video_inspect_transcribe_required.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: inspect-stt-001" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/talk.mp4","frames":false,"transcribe":true,"duration_seconds":480,"language_code":"en","segmentation":{"mode":"sentence"}}'Longer than 10 minutes
The inspect source limit is 1,800 s, but the hint caps at 600. For a 25-minute recording, detach the audio in ranges with audio detach (each range up to 900 s) and transcribe those files with the STT endpoint, or run inspect and check the live reserve in the catalog.
Silent clips
A clip with no audio track returns inspect_source_has_no_audio. Run a frames: false inspect first and read probe.has_audio. Inspect returns facts, stills and text. For semantic questions about what happens in the video, the docs point to separate tools that are not on every environment.
Budget
For 100 clips of 8 minutes the STT part is 100 x $0.08 = $8.00, plus compute for each inspect. If you only need the transcript, send frames: false so the job does not spend time on stills. Keep transcripts for later queries; the result includes text, words[] and optional sentence segments[]. Every media job follows the same lifecycle: submit with an Idempotency-Key, receive a job, poll GET /v1/jobs/:id/status until it is ready, then read GET /v1/jobs/:id/result. A retry with the same key does not queue a second job, so a network error during submit never doubles a charge.
Sources
Related posts
More in Media tools
- Video upscale API: a 12-second clip is $0.108, the cap is 30 seconds
Sume Video Upscale 1.0 bills $0.009 per input second, up to 30 seconds ($0.27). Factor 1.1 to 4, three tiers, and a 5-second default reserve.
- Vidu Q4 540p draft to 1080p: Sume video upscale at 2x costs $0.09
A 10-second 960x540 draft upscaled at the default factor 2 gives 1920x1080 for $0.09 on Sume. Sume does not list Vidu; the clip only needs a public HTTPS URL.
- Voice greeting on a still photo: TTS first, then H3 Max lip sync
To make a still photo speak, make 5 to 14.8 seconds of TTS audio, then call MiniMax H3 Max lip sync. TTS costs $0.0475 per 1,000 characters.
- Voiceover plus music bed: TTS and Lyria cost, duck_db in Timeline
A 700-character voiceover ($0.03325) and one Lyria bed ($0.125) cost $0.15825 before the render. Timeline duck_db (0 to 20) lowers the bed under speech.
Written by Sume