Pick the cleanest last frame: sample 24 stills before chaining
Before chaining AI clips, sample up to 24 stills with Sume video frames (fps up to 2) and choose the best one as the next first_frame, not just the last.

When you chain AI clips, the very last frame is often a bad starting still: it can be blurred, mid-gesture or closing on a transition. Sume's video frames API lets you sample up to 24 stills from a clip, so you can pick the cleanest one and send it as the next clip's first_frame.
This replaces the 16 keyframes Luma lists for Ray 3.2, which Sume does not run, with a manual keyframe choice between clips.
Sample, do not just grab the end
Video frames takes one workspace clip and either at[] (1 to 24 explicit seconds) or fps (above 0 and at most 2). It returns durable media.sume.com image artifacts at the source size, as jpeg (default) or png, with an optional max_edge of 16 to 2160. The source must be 300 seconds or shorter.
With fps, Sume uses mid-bin samples: 0.5/fps, 1.5/fps and so on, with a cap of 24 frames. On a 10-second clip, fps: 2 gives 20 stills, and fps: 0.8 over a 30-second clip gives 24 stills with the last at 29.375 seconds.
| Clip length | fps | Frames | Last sample at |
|---|---|---|---|
| 5 s | 2 | 10 | 4.75 s |
| 10 s | 2 | 20 | 9.75 s |
| 12 s | 2 | 24 | 11.75 s |
| 30 s | 0.8 | 24 | 29.375 s |
The request
A submit always returns 202, because the route pins async mode, so do not send mode: sync. Poll GET /v1/video-frames/:id until resource_status is ready. A frame that fails to extract has a url of null and does not fail the job.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: frames-clip-07" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/clip07.mp4",
"fps":0.8,"format":"png"}'What to look for in the stills
Choose a frame with the subject in a neutral pose, open eyes, no motion blur and the framing you want the next shot to start with. Prefer png for the pick, since it is lossless, and the frame keeps the source size unless you set max_edge.
If you ask for an explicit at of 30 on a 30-second clip, the worker fails with frame_time_out_of_range, because t must be at least 0 and below the duration. Ask for 29.9 instead.
Feed it forward
Send the chosen still to the next POST /v1/videos as a frame_images entry with frame_type: first_frame, on a model whose supported_frame_images lists it. For an extra anchor, also send a last_frame image if the model supports that type. Each frame extraction is billed by its Modal compute (container seconds times Modal list times 1.25, plus the platform fee), and the bill is never more than the hold.
A last-frame chain accumulates drift: every hop is a fresh generation from a still. Keep chains short, and rejoin at hard cuts in Timeline.
Using video inspect instead
If you only need a quick look, video inspect returns a probe plus 8 mid-bin stills by default, at a default max_edge of 768. That is enough to judge a clip but too small for a first frame. Use video frames, which keeps the source size, when you pick a still for the next job.
Video frames allows 1 to 24 values in at[] or an fps up to 2, and the source must be 300 seconds or shorter. A 30-second AI clip is well within that.
A selection routine
- Sample the last 3 seconds with an explicit
atlist, for example 27, 27.5, 28, 28.5, 29, 29.5 on a 30-second clip. - Reject frames with motion blur, closed eyes or a half-turned head.
- Prefer a frame the model could plausibly continue from, with space around the subject.
- Send the winner as
first_frameand keep the same prompt style for the next clip.
Cost and failure notes
The extract reserves its Modal compute ceiling at submit and bills the actual container time times the Modal list times 1.25, plus the platform fee, never more than the hold. Because a failed instant returns a null url without failing the job, check every frame URL before you pick. For a clip over 90 seconds, warnings[] can include low_confidence_long_video; AI clips under 30 seconds do not hit it.
Sources
Related posts
More in Media tools
- Pinterest video ads take H.264 or H.265: do you need H.265?
Pinterest video ads accept H.264 or H.265 in MP4, MOV or M4V, so an H.264 clip is fine. Sume's exact trim uses libx264; confirm any other output with a probe.
- Podcast quote clips end abruptly: STT boundary_lead_ms tail padding
Sume STT segmentation boundary_lead_ms (0 to 500 ms, default 70) sets how long a sentence's tail runs before the cut. Tune it, then cut with timeline audio.
- Product spec sheet to video: Wan 3.0 footage, exact specs via compose
Make a spec-sheet video without a model re-typing your numbers: generate the footage with Wan 3.0, then put your own spec card on screen with Timeline compose.
- Product turntable ad: which Sume video models take two frames
For a product turn, give a video model the front and back photo as first and last frames. Six Sume rows take both; Grok Imagine takes a first frame only.
Written by Sume