Veo 3.1 first and last frame: how the API works

Veo 3.1 takes a first image and a last image and generates the motion between them. Here is Google's request shape and where the same input works on Sume.

4 min readSume
All posts

Veo 3.1 can start from a first image and end on a last image: you pass the first as image and the last as lastFrame, and the model generates the motion between them. Google marks this as available for Veo 3.1 models only, and lastFrame must be used together with image. Sume's video catalog does not currently list a Veo model. On Sume, the same idea is frame_images with first_frame and last_frame entries, on the models below.

Google's parameters are from its Veo page; Sume's are from the Video generation docs, read 2026-09-29.

What does Google's Veo 3.1 last-frame request need?

Google's parameter table lists image (an initial image to animate) and lastFrame (the final image for an interpolation video to transition, only with image). Aspect ratio is 16:9 (default) or 9:16, duration is 4, 6 or 8 seconds, and resolution defaults to 720p. Google calls this interpolation; personGeneration for it is allow_adult only.

How do I send first and last frames on Sume?

Use frame_images on POST /v1/videos. Each entry needs a frame_type of first_frame or last_frame. A last frame without a first frame is refused. If you also send input_references, frame_images wins and the request is image-to-video.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: frames-001" \
  -d '{
    "model": "seedance-2.5",
    "prompt": "The door opens and morning light fills the room.",
    "frame_images": [
      {"type":"image_url","image_url":{"url":"https://example.com/closed.png"},"frame_type":"first_frame"},
      {"type":"image_url","image_url":{"url":"https://example.com/open.png"},"frame_type":"last_frame"}
    ],
    "duration": 6
  }'

Which Sume models accept a last frame?

From the catalog code, the models that list both frame types are below. grok-imagine-video-1.5 takes a first frame only, and it needs one.

Frame support in the Sume video catalog, read 2026-09-29.
ModelFrame typesClip length
seedance-2.5first, last4 to 30 s
seedance-2, seedance-2-fast, seedance-2-minifirst, last4 to 15 s
kling-3first, last4 to 15 s
wan-3.0first, last2 to 30 s
minimax-h3, minimax-h3-maxfirst, last5 to 15 s
gemini-omni-flash-1.1first, last3 to 10 s
grok-imagine-video-1.5first only4 to 15 s

What should the two frames look like?

Keep both images the same size and framing, and use public HTTPS URLs. The model invents everything between them, so if the two frames differ a lot, ask for one simple move in the prompt. Check the first frames of the result before you use it; if the ending drifts, shorten the duration.

Sources

Related posts

More in Models

All Models posts

Written by Sume