Transcript from a video on Sume: inspect transcribe or detach first?
Sume can transcribe a video through video inspect (1,800 s limit, compute plus $0.01 per minute) or from a detached 16 kHz mono wav. How to choose.

Use video inspect with transcribe true when you also want probe data or stills from a clip of 30 minutes or less. Use audio detach first when the video is longer, when you want a reusable audio file, or when you want to control the exact range and format. Both end at Sume STT 1.0, which is listed at $0.01 per audio minute.
The two paths
Video inspect runs STT on the audio of a clip. It reserves its Modal compute ceiling and captures the actual container seconds times the Modal list times 1.25, plus the platform fee, never more than the hold. The per-minute STT rate is added to the reservation. The source limit is 1,800 seconds. You can set language_code, a duration hint (up to 600 seconds), and segmentation.mode sentence for caption-shaped segments.
Audio detach compiles ffmpeg on the worker and returns a new wav or mp3 for $0.01 per job, with no provider inference. Source up to 1,800 seconds and output up to 900 seconds. The wav (pcm_s16le) is the format that speech-to-text uses, and 16 kHz mono is the STT shape.
| Question | Video inspect with transcribe | Audio detach then STT |
|---|---|---|
| Source limit | 1,800 s | 1,800 s, output 900 s per job |
| Extra output | Probe, up to 24 stills per call | A durable audio file |
| Fees besides STT | Compute capture plus platform fee | $0.01 per detach job |
| Language and segments | language_code, sentence segmentation | Set on the STT step |
| Silent video | inspect_source_has_no_audio | detach_source_has_no_audio |
How to decide
If you are making captions or a quick summary of a short clip, inspect is one call. If you are building a pipeline (dubbing, lyric files, an archive), detach once and reuse the audio: later steps such as Timeline audio split read the detached file, and you never decode the video again.
For a 47-minute video neither path works in one call as written: inspect refuses a source over 1,800 seconds. Detach it in two ranges of 900 seconds or fewer, then transcribe each wav.
Shared pitfalls
Both paths need a source on your workspace's media host. Import the file first. Both refuse a silent source, so probe has_audio once. And both bill per minute of audio only at the STT step, so trimming silence off the start and end before the transcript is a real saving at volume: 1,000 files with 30 seconds of dead air each is 500 minutes, or $5.00.
Cost at volume
At 500 clips a month of 10 minutes each, the STT minutes alone are 5,000 x $0.01 = $50. Detach adds 500 x $0.01 = $5. Inspect adds compute and a platform fee on each call, so measure one call before you extrapolate. Whichever path you use, read the live rates in GET /v1/catalog.
Sources
Related posts
More in Developers
- TTS word timestamps: timestamps.words and sentence segmentation
Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.
- Turn a roleplay debrief into an avatar feedback clip in Python
Take the written debrief from a roleplay or survey session and render it as a 16:9 Sume avatar clip with a retry-safe key, a 12 to 168 word check and polling.
- 12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4
A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.
- TypeScript 7.0.2 with @sume-com/sdk 0.2.0: the tsconfig that compiles
We compiled a Sume SDK 0.2.0 job-status call with TypeScript 7.0.2. It needs nodenext resolution, types set to node and a module-type package.json.
Written by Sume