video_inspect is not free: a $0.0899 Modal hold, then actual seconds
video_inspect now holds a $0.0899 Modal ceiling and captures its own container seconds at Modal list x 1.25 plus the 5.5% fee. Transcripts add $0.01 a minute.

video_inspect used to be described as free. It is not any more. Each call reserves a Modal ceiling of **$0.0899** (8 cores, 4 GiB, 600 seconds), then captures what the container actually used: seconds x Modal list x 1.25, plus the 5.5% platform fee. A transcript adds $0.01 per audio minute.
What changed
Older articles, including some on this blog, say probes and stills cost nothing. The current docs for video_inspect, video_frames and reference_ingest say each reserves a Modal ceiling and captures its own container seconds. The probe, frame sample and still extraction run on the shared media_ffmpeg shape.
The ceiling is a hold, not a bill: the unused part is released when the job ends.
How the $0.0899 is built
Modal's pricing page lists $0.0000131 per physical core-second and $0.00000222 per GiB-second. The media_ffmpeg shape is 8 cores and 4 GiB for up to 600 seconds.
| Step | Arithmetic | Result |
|---|---|---|
| Cores | 8 x $0.0000131 per second | $0.0001048 |
| Memory | 4 x $0.00000222 per second | $0.00000888 |
| 600 seconds | $0.00011368 x 600 | $0.068208 |
| x1.25 | $0.068208 x 1.25 | $0.08526 |
| +5.5% fee | $0.08526 x 1.055 | $0.08995 |
The transcript is separate
If you ask for a transcript, STT is a second line: $0.01 per audio minute. With no duration hint, one minute is reserved. The maximum hint is 600 seconds, so the transcript hold tops out at $0.10 on top of the Modal ceiling. engine: scribe_v2 makes no Modal call and keeps its own STT hold.
What to budget
For a probe-only call, $0.0899 is the most it can hold, and the capture is the seconds it actually ran. For a ten-minute clip with a transcript, plan on at most the Modal ceiling plus $0.10. Read the real number from the usage row afterwards.
Related jobs that moved with it
The same change covers video_frames and reference_ingest. All three ride the media_ffmpeg shape, so they share the same hold. Their docs describe the charge the same way: each reserves a Modal ceiling and captures its own container seconds x Modal list x 1.25 plus the platform fee.
reference_ingest adds speech-to-text only when there is speech, which keeps the transcript line off silent footage. The Modal line is still there, because the container still ran.
What to change in your own estimates
If you wrote a cost model that treated inspect, frames and ingest as free, add a per-call line. Use the ceiling as the upper bound and measure the real capture from usage rows for the typical case.
If you only inspect a few clips an hour, the line is negligible. If a pipeline fans out across hundreds of clips, it is no longer rounding error, and the 5.5% fee applies to it like any other spend.
Sources
Related posts
More in Media tools
- Video inspect 400 video_inspect_transcribe_required: the fix
Sending language_code, segmentation or duration_seconds to Sume video inspect without transcribe true returns a 400. Why, and the request that works.
- Video trim limits on Sume: 1800 s source, 900 s output, error codes
Video trim takes a source up to 1800 seconds and cuts at most 900 seconds, at least 0.2. Here are the limits, the clamp warning and the stable error codes.
- Which field holds the file URL in a Sume media job result
Sume media jobs return different result keys: audio_url for detach, video_url for trim and filter, frames for stills. A field map for no-code steps.
- Korean caption styles on Sume: black-outline, clip-wipe and the rest
For Korean speech on Sume, pick a Hangul caption style: black-outline is the safe default, clip-wipe reads best small. All six, and why slam shows tofu.
Written by Sume