Use a frame from your footage as a reference for AI video
Pull a still at a time you pick with video frames (unbilled), then pass it as image_url or reference_image_urls to a Sume video model. Steps and limits.

Extract the frame with POST /v1/video-frames, naming the time in at[], then pass the returned image URL to a video model: as image_url for a start frame, or in reference_image_urls (up to 10) for a reference. Adobe's Premiere 26.5 notes say its Generative Media Tool lets editors use existing frames as visual references; on Sume you do the same in two API calls.
Details are from Video frames and Video Router, read 2026-09-30.
How do I pull the frame?
Video frames takes one media.sume.com clip and exactly one of at[] (1 to 24 times in seconds) or fps. Optional format is jpeg (default) or png, and max_edge is 16 to 2160; omit it to keep the source size. The call is always 202 async, and the call is unbilled. Each time must be within the clip, or the worker fails with frame_time_out_of_range.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ref-frame-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"at": [4.2],
"format": "png"
}'How do I hand the frame to a video model?
When resource_status is ready, frames[{t,url,width,height}] hold durable image URLs. On gemini-omni-flash-1.1 the request shape picks the capability, as the table shows.
| You want | Send |
|---|---|
| Frame as the opening shot | image_url (plus optional end_image_url) |
| Frame as a look or subject reference | reference_image_urls (up to 10), addressed as <IMAGE_REF_0> in the prompt |
Which frame should I choose?
Pick a frame where the subject is sharp and the framing is what you want the new clip to start from. Not only the last frame: at[] takes any time, so you can pull several candidates in one call and compare. For the last-frame case, see get the last frame of a video to chain clips.
What are the limits?
The source clip can be up to 300 seconds, at most 24 frames per call, and it must be a Sume-hosted artifact (import it first with POST /v1/media-imports). The video model's own limits apply to the new clip: gemini-omni-flash-1.1 is 3 to 10 seconds at 16:9 or 9:16. A reference guides the generation; it does not guarantee the new clip matches the original footage.
Sources
Related posts
More in Use cases
- Remove background from video with AI: what Sume covers
Sume has no video background-removal endpoint: RMBG 1.0 takes a still image. What you can do instead is a prompt edit with video_to_video, or a crop.
- Replace video audio with an AI voice (and the lip-sync catch)
Sume has no one-call audio swap: detach the audio, transcribe it, make a new voice with TTS, and lay it on a Timeline. Lips will not re-sync to the new voice.
- AI ad resizer: one image, several aspect ratios, one API
Sume has no resizer button. Send the same source image to POST /v1/images once per aspect_ratio (1:1, 4:5, 9:16, 16:9) and recompose each ad slot.
- Seedance 2.5 draft mode: a 480p-then-1080p loop on Sume
Higgsfield added Seedance 2.5 Draft Mode. On Sume there is no draft model: render at resolution 480p, pick a take, then re-request at 1080p.
Written by Sume