Video frames fps on Sume: mid-bin sampling and the 24-frame cap
Sume video frames with fps 0.5 samples at 1 s, 3 s, 5 s and so on, capped at 24 frames per call, at source size. Use at[] for exact instants.

With fps: 0.5, Sume video frames samples mid-bin instants, 0.5/fps, 1.5/fps and so on, so the stills land at 1, 3, 5 seconds, not at 0. The call returns at most 24 frames, fps must be above 0 and no more than 2, and the clip must be 300 seconds or shorter. For exact instants send at[] instead. Source: Video frames, read 2026-10-06.
How do at[] and fps differ?
Send only one of them.
| Program | Values | Result |
|---|---|---|
at[] | 1 to 24 seconds, each at least 0 and below the duration | One still per instant; out of range gives frame_time_out_of_range |
fps | Above 0, at most 2 | Mid-bin samples, limit 24 frames |
format | jpeg (default) or png | Lossless with png |
max_edge | 16 to 2160 | Omit it and frames keep the source size |
When do I use video frames and not video inspect?
Video inspect samples 8 mid-bin stills with a default max_edge of 768. Video frames does not clamp unless you send max_edge, so use it when you need the source size, such as an OCR pass or a check of a first frame. A submit always returns 202, because the job is pinned to async, and you read GET /v1/video-frames/:id until resource_status is ready.
- A failed instant has
url: nulland does not fail the job. - Over 90 seconds,
warnings[]can carrylow_confidence_long_video. - Billing is by the job's Modal compute, never more than the reserved hold.
Sources
Related posts
More in Media tools
- Video inspect on Sume: 8 stills at 768 px and the fast seek option
Sume video inspect returns 8 mid-bin stills at a 768 px long edge by default, up to 24 stills, and a fast seek mode that can land one GOP early.
- AI video upscaler API fields: factor 1.1 to 4 and fast, standard, pro
The Sume Video Upscale request takes upscale_factor from 1.1 to 4, duration_seconds from 1 to 30 and an enhancement_tier. What they do and cost.
- Voiceover longer than your clips: Timeline pads, loops, 0.5 s rule
If the voiceover outruns your clips, Sume Timeline still renders to audio.duration_seconds: short sources pad or loop, and coverage may stop 0.5 s early.
- Why a voiceover on an AI video clip never lip-syncs
Video models do not lip-sync to TTS or a later voiceover. Per Sume's docs a talking face comes from Fabric, H3 Max Lip Sync or an avatar talking-video job.
Written by Sume