Chain two AI clips: last frame in, then first_frame out
To continue a clip on Sume, extract its last frame with video_frames, then send it as first_frame on the next job. It fixes pose and position, not sound.

To continue a clip, extract its last frame with POST /v1/video-frames, then send that still as a first_frame entry in frame_images on the next video job. The second clip opens where the first one ended, so position, pose and set carry over. Lighting, voices and background sound do not, because each job generates its own.
Use this when the target length is longer than one job allows, or when you want to redo only the second half of a piece.
The steps
Frame extraction works on one clip hosted on media.sume.com. If your result URL is not already a media.sume.com artifact, import it first with POST /v1/media-imports. Pick a time just before the end, such as 14.8 for a 15-second clip, because the very last instant can be a blend frame.
| Step | Call | Result |
|---|---|---|
| 1. Make clip one | POST /v1/videos | A completed job |
| 2. Extract a frame | POST /v1/video-frames with at: [14.8] | A durable image URL |
| 3. Make clip two | POST /v1/videos with frame_images, frame_type: first_frame | A clip that opens on that frame |
| 4. Join | Timeline render | One finished file |
A request for clip two
Only ids that support image-to-video take the first frame. Check supported_frame_images on the catalog row. The body looks like this.
{
"model": "wan-3.0",
"prompt": "The camera keeps moving forward through the doorway",
"duration": 10,
"resolution": "720p",
"frame_images": [
{
"type": "image_url",
"image_url": {"url": "https://media.sume.com/artifacts/your-frame.jpg"},
"frame_type": "first_frame"
}
]
}Where it breaks
Drift is real. A face seen from a new angle can change, a logo can lose a letter, and the generated audio will start fresh at the cut. Hide the join with a cut or a short fade in the timeline instead of trying to make it invisible. If one continuous shot matters more than price, use an id that goes to 30 seconds in a single job, seedance-2.5 or wan-3.0.
Checks before the join
Look at the two ends before you cut them together.
Frame extraction is billed by its own compute, separate from the video job. It is small next to the generation, but include it in a batch estimate.
Write the prompt for clip two as a continuation, not a fresh scene. Name what the camera does next and what stays fixed, because the model sees the still and your text, and nothing of the motion that led up to it.
- Compare the last frame of clip one with the first of clip two.
- Listen to both audio tracks at the cut.
- Use a short fade if the lighting differs.
- Keep the same resolution and aspect ratio on both jobs so the render does not rescale.
Sources
Related posts
More in Use cases
- Homebrew beer label art with an API: the name and ABV added in code
Generate square label art on Sume with no text, then draw the beer name, style and ABV in code so every batch gets exact numbers on the same artwork.
- Hotel room tour clip from 10 photos with Omni Flash 1.1
Gemini Omni Flash 1.1 on Sume takes up to 10 reference images and 3 short clips. An 8-second room tour costs $1.00 at 720p, $1.50 at 1080p and $3.00 at 4K.
- How long can an AI video be? Shorts 3 min, TikTok ads 10 min
YouTube allows Shorts up to 3 minutes and TikTok non-Spark ads up to 10. A Sume clip is at most 30 s, but Timeline renders up to 1800 s from several clips.
- How to split a training lesson into avatar clips under 60 seconds
Sume refuses avatar videos over 60 seconds. Split a 6-minute lesson by words, render each part, and stitch with Timeline 1.0. Word counts and cost by tier.
Written by Sume