AI video transition generator: bridge one shot to the next

An AI transition is a generated clip that starts on the last frame of one shot and ends on the first frame of the next. How to make one with Sume.

5 min readSume
All posts

An AI video transition generator makes a short new clip that carries one shot into the next: the clip starts on the last frame of shot A, ends on the first frame of shot B, and a video model generates the motion between them from your prompt. You pull the two frames, send them as a first and a last frame, and cut the new clip in between. A plain fade, wipe, or slide needs no generation at all.

The steps use Sume's Video frames, Video generation, and Timeline 1.0 docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's API code.

How do I generate a transition between two clips?

Three steps, with both shots finished:

  • Pull the two frames. If both shots are Sume outputs, POST /v1/video-frames returns exact stills at the times you name as durable media.sume.com images, unbilled. Ask for 0 on shot B, and for a time just before the end of shot A, because every time must satisfy 0 <= t < duration; extracting the last frame shows the arithmetic. For footage shot elsewhere, export the two stills from your editor and host them at public HTTPS URLs.
  • Generate the bridge. Send both stills to POST /v1/videos in frame_images: shot A's last frame as first_frame and shot B's first frame as last_frame, with a prompt that describes the move between them. Set aspect_ratio to your shots' shape: in current code, a Seedance request without one asks the model for 9:16.
  • Cut it in. Place the bridge between shot A and shot B in your editor, or in a Timeline 1.0 render when all three clips are in your workspace on media.sume.com.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: bridge-a-to-b-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "One continuous move: the camera pushes through the cafe window and comes out over the beach at sunset",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://media.sume.com/artifacts/artf_demo/a-last.png" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://media.sume.com/artifacts/artf_demo/b-first.png" }, "frame_type": "last_frame" }
    ],
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 4
  }'

What should an AI transition prompt say?

No field sets the kind of transition, so it all goes in prompt. Sume's docs suggest details about motion, camera angles, lighting, and scene composition. Describe one continuous move that could carry the first image into the second, for example:

  • Push through: "The camera pushes through the cafe window and comes out over the beach at sunset."
  • Whip pan: "A fast whip pan to the right blurs the street, and the frame settles on the mountain lake."
  • Morph: "The steam from the coffee cup swirls upward and becomes the clouds over the valley."
  • Zoom: "The camera zooms into the phone screen until its picture fills the frame."

How long is a generated transition?

At least as long as the model's shortest clip: duration must be a length the model lists, and in current code any other value is refused. The bridge adds that time to your edit, so pick the model by its minimum and by whether it takes a last frame.

From Video generation, Video Router, the Sume API reference, and the catalog behind GET /v1/videos/models, read 2026-09-28.
Model idShortest clipTakes a `last_frame`
wan-3.02 sYes
gemini-omni-flash-1.13 sYes
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini4 sYes
kling-34 sYes
minimax-h3, minimax-h3-max5 sYes
grok-imagine-video-1.54 sNo

Is an AI transition the same as a match cut?

No. A match cut is an editing term for a hard cut between two shots whose shapes or motion line up, with no new frames. A generated transition adds new frames between the shots instead. If your two frames already line up, a plain cut may be all you need.

When is a plain transition enough?

When you only need to soften the cut. A Timeline 1.0 render can put a fade, wipe, slide, or dissolve of up to 1 second between two clips with no model inference, as Video transitions API explains. A short fade into and out of a generated bridge combines the two.

What are the limits?

  • Video frames and Timeline read only your workspace's media.sume.com files. Sume mirrors generated clips to its own media URLs, so earlier Sume outputs qualify: once a job is completed, read the clip's URL from GET /v1/jobs/{id}/result.
  • Frame images for generation must be public HTTPS URLs. A last_frame needs a first_frame in current code, and grok-imagine-video-1.5 takes no last frame.
  • The model invents the frames between the two ends. Nothing promises a smooth or exact match, so watch the bridge before you cut it in.
  • A Timeline render takes its sound from the audio spine and the optional soundtrack. In current code each clip's own audio is dropped.
  • Every bridge is its own generation, billed at the rates the model lists in pricing_skus on GET /v1/videos/models.

Sources

Related posts

More in Models

All Models posts

Written by Sume