Modal holds: video_frames $0.0899, hyperframes_preview $0.1455
Sume sizes Modal holds per job family: media_ffmpeg at $0.0899, hyperframes_preview at $0.1455, hypit verbs at $0.0121. Which jobs share which ceiling, and why.

Sume holds money against a fixed Modal shape before a media job runs, then captures actual container seconds. video_inspect, video_frames, video_frame_sample and reference_ingest share a **$0.0899** ceiling. hyperframes_preview holds **$0.1455**. Hypit verbs (probe, boundaries, tile, cut) hold only **$0.0121**.
The ceiling table
Each hold is the shape's Modal list cost over its maximum seconds, times 1.25, plus the 5.5% fee. Modal lists $0.0000131 per core-second and $0.00000222 per GiB-second.
| Ceiling | Jobs | Shape x max seconds | Hold |
|---|---|---|---|
| media_ffmpeg | video_inspect, video_frames, video_frame_sample, reference_ingest | 8 cores / 4 GiB x 600 s | $0.0899 (+ STT minutes) |
| hyperframes_preview | hyperframes_preview | 8 cores / 8 GiB x 900 s | $0.1455 |
| hypit_verb | probe, boundaries, tile, cut | 2 cores / 2 GiB x 300 s | $0.0121 |
Hold versus charge
A hold only keeps your balance from going negative mid-job. The charge is the container seconds the job really used, priced at the same rate. A job that ends in 20 seconds on the media_ffmpeg shape captures 20 seconds of container time, not 600, and releases the rest.
If a job ever ran past its ceiling, the capture is clipped to the hold. The usage row records this with price_book_clipped_to_hold: 1. You never pay more than the ceiling you were quoted.
Why hyperframes_preview costs more
The preview shape doubles the memory to 8 GiB and runs up to 900 seconds instead of 600. That is the entire difference: more GiB-seconds and more seconds. The per-second rates are identical.
Using it in practice
If you are building a pipeline that calls inspect and frame sampling many times, the ceiling is what must be available in your balance per in-flight job. Ten parallel video_inspect calls need about $0.90 of headroom, not ten times the typical charge. Keep a balance above that if your account runs concurrent media jobs.
Why the shapes are different sizes
The shapes reflect what the work needs. Probing a file and pulling stills is mostly I/O and ffmpeg, so 8 cores and 4 GiB is sized for a 600-second ceiling. Hypit verbs are lighter, which is why they hold 2 cores and 2 GiB for 300 seconds.
Because every ceiling uses the same Modal rates, comparing them is just comparing core-seconds and GiB-seconds. You do not need a separate price list per job.
Reading the hold on a usage row
When a hold is released or captured, the row shows what happened. If the capture is smaller than the hold, the remainder was released. If you see a clipped flag, the job tried to use more than its ceiling.
None of these jobs has a surprise multiplier. Everything is list x 1.25 plus the fee, bounded by the table above.
Sources
Related posts
More in Media tools
- Video inspect stills come back 432x768: set max_edge 1920
Video inspect clamps stills to a 768 long edge by default, so a 1080x1920 clip returns 432x768 frames. Set max_edge up to 2160, or use video frames.
- Video inspect frames: at and fps together return a 400 conflict
Sume video inspect frames takes either at[] or fps, never both. The conflict and required codes, the 24-still cap, and how to sample a clip evenly.
- video-inspect 400: frames at and fps together on a Shorts hook
Send frames.at or frames.fps to Sume video-inspect, never both: both returns 400 video_inspect_frames_program_conflict. Pick at[] for a Shorts hook check.
- video-inspect frames needs at or fps: probe a Short with frames:false
video_inspect_frames_program_required means frames is an object with no at or fps. Omit frames for 8 stills, or pass frames:false to only probe a Short upload.
Written by Sume