Chain two AI clips: last frame in, then first_frame out

To continue a clip on Sume, extract its last frame with video_frames, then send it as first_frame on the next job. It fixes pose and position, not sound.

5 min readSume
All posts

To continue a clip, extract its last frame with POST /v1/video-frames, then send that still as a first_frame entry in frame_images on the next video job. The second clip opens where the first one ended, so position, pose and set carry over. Lighting, voices and background sound do not, because each job generates its own.

Use this when the target length is longer than one job allows, or when you want to redo only the second half of a piece.

The steps

Frame extraction works on one clip hosted on media.sume.com. If your result URL is not already a media.sume.com artifact, import it first with POST /v1/media-imports. Pick a time just before the end, such as 14.8 for a 15-second clip, because the very last instant can be a blend frame.

Last-frame handoff between video jobs (read 2026-10-06 from the Sume docs)
StepCallResult
1. Make clip onePOST /v1/videosA completed job
2. Extract a framePOST /v1/video-frames with at: [14.8]A durable image URL
3. Make clip twoPOST /v1/videos with frame_images, frame_type: first_frameA clip that opens on that frame
4. JoinTimeline renderOne finished file

A request for clip two

Only ids that support image-to-video take the first frame. Check supported_frame_images on the catalog row. The body looks like this.

{
  "model": "wan-3.0",
  "prompt": "The camera keeps moving forward through the doorway",
  "duration": 10,
  "resolution": "720p",
  "frame_images": [
    {
      "type": "image_url",
      "image_url": {"url": "https://media.sume.com/artifacts/your-frame.jpg"},
      "frame_type": "first_frame"
    }
  ]
}

Where it breaks

Drift is real. A face seen from a new angle can change, a logo can lose a letter, and the generated audio will start fresh at the cut. Hide the join with a cut or a short fade in the timeline instead of trying to make it invisible. If one continuous shot matters more than price, use an id that goes to 30 seconds in a single job, seedance-2.5 or wan-3.0.

Checks before the join

Look at the two ends before you cut them together.

Frame extraction is billed by its own compute, separate from the video job. It is small next to the generation, but include it in a batch estimate.

Write the prompt for clip two as a continuation, not a fresh scene. Name what the camera does next and what stays fixed, because the model sees the still and your text, and nothing of the motion that led up to it.

  • Compare the last frame of clip one with the first of clip two.
  • Listen to both audio tracks at the cut.
  • Use a short fade if the lighting differs.
  • Keep the same resolution and aspect ratio on both jobs so the render does not rescale.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume