Pick a thumbnail from a clip on Sume: video-frames or video-inspect
Use video-frames when you need a still at exact times and source size. Use video-inspect for a quick look: 8 stills by default, with a fast seek mode.
For a thumbnail at times you choose, use POST /v1/video-frames with at[]: it returns durable stills at the source size. For a quick look at what is in a clip, use POST /v1/video-inspect, which returns a probe plus eight mid-bin stills by default at a 768-pixel long edge. Both bill by Modal compute, not at a flat rate.
video-frames: exact times, source size
Send video_url and exactly one of at[] or fps. Options are format (jpeg by default, or png) and max_edge from 16 to 2160. Without max_edge, frames keep the source frame size. A submit always returns 202; poll GET /v1/video-frames/:id until resource_status is ready.
The result is frames[{t,url,width,height}] as durable artf_ images. If extraction fails at one instant, that frame has url: null and the job does not fail.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-frames-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"at": [0, 2.5]
}'video-inspect: probe, a contact sheet and optional speech
Inspect reads one clip and returns probe facts, stills, and an optional transcript. The default mode is sync and waits up to 30 seconds. If you do not send frames, you get eight mid-bin stills, or one per second if the clip is shorter than eight seconds. Send frames: false for the probe alone, with no stills.
The seek setting chooses how each still is found: precise (default) decodes to the accurate instant, and fast moves to the keyframe at or before it. Fast can be earlier by up to one GOP, roughly 0 to 5 seconds on typical sources, but never later.
Which one
| You need | Use | Notes |
|---|---|---|
| A thumbnail at 2.5 s, at source size | video-frames with at | Up to 24 stills a call; max_edge 16 to 2160 |
| A look at the whole clip | video-inspect, default frames | 8 stills, max_edge default 768 |
| Only duration, size or whether there is audio | video-inspect with frames: false | Check probe.has_audio before detach |
| A fast skim of a long clip | video-inspect with seek: fast | Stills can be one GOP early |
| Words with timings | video-inspect with transcribe: true | $0.01 per audio minute |
Limits and costs
Both accept sources up to 1800 seconds and 24 stills for each call. The sources must be this workspace's media.sume.com files, so import a public clip first. The bill is container seconds times the Modal list price times 1.25, plus the platform fee, and it never exceeds the hold made at submit.
Choosing the picture
Take a wide spread of times first, then look. A still at 0 seconds is often a fade or a logo. Ask for several times and pick one, rather than trusting a single guess.
Related posts
More in Media tools
- Pick the cleanest last frame: sample 24 stills before chaining
Before chaining AI clips, sample up to 24 stills with Sume video frames (fps up to 2) and choose the best one as the next first_frame, not just the last.
- Pinterest video ads take H.264 or H.265: do you need H.265?
Pinterest video ads accept H.264 or H.265 in MP4, MOV or M4V, so an H.264 clip is fine. Sume's exact trim uses libx264; confirm any other output with a probe.
- Podcast quote clips end abruptly: STT boundary_lead_ms tail padding
Sume STT segmentation boundary_lead_ms (0 to 500 ms, default 70) sets how long a sentence's tail runs before the cut. Tune it, then cut with timeline audio.
- Product spec sheet to video: Wan 3.0 footage, exact specs via compose
Make a spec-sheet video without a model re-typing your numbers: generate the footage with Wan 3.0, then put your own spec card on screen with Timeline compose.
Written by Sume