CTA end card for an AI video ad: use last_frame on Sume

End an ad clip on your CTA card by sending it as a last_frame image on /v1/videos. Which models accept it, which do not, and the Timeline alternative.

6 min readSume
All posts

To make a video ad end on your call-to-action card, send the card as a frame_images entry with frame_type last_frame in a POST /v1/videos body. Sume accepts a first frame, a last frame or both, but only on models whose catalog entry lists last_frame in supported_frame_images. If you need the card to appear exactly, unchanged and for a fixed number of seconds, use Timeline 1.0 instead and append the still as a static hold.

Which models take a last frame

The Video Router catalog marks an end-frame capability per model, and the /v1/videos descriptor turns that into supported_frame_images. The table is taken from the capability records in the Sume repository and should be confirmed against GET /v1/videos/models for the model you use.

End-frame support per video model in the Sume catalog code, read 2026-10-07
Model idLast frameDuration range (s)
seedance-2.5Yes4 to 30
seedance-2, seedance-2-mini, seedance-2-fastYes4 to 15
kling-3Yes4 to 15
wan-3.0Yes2 to 30
minimax-h3, minimax-h3-maxYes5 to 15
gemini-omni-flash-1.1Yes3 to 10
grok-imagine-video-1.5No4 to 15
h3-max-recastNo (edit-only model)5 to 30 source

The request

Each frame is an object with type image_url, an image_url object holding a public URL, and a frame_type. If you also want the clip to start from a product still, add a second entry with frame_type first_frame. The API infers image-to-video from the presence of frame_images, so no mode field is needed.

The model then generates motion that ends on, or near, your last frame. The result is a clip that arrives at the card, not a clip with a card pasted on. You should still expect small differences in how closely the final frame matches your image, so check the last second of each output before you ship it.

curl -sS -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: mug-endcard-v1" \
  -d '{
    "model": "seedance-2",
    "prompt": "Hand sets a mug down, camera pushes in",
    "duration": 6,
    "aspect_ratio": "9:16",
    "frame_images": [
      {"type":"image_url","image_url":{"url":"https://example.com/end-card.png"},"frame_type":"last_frame"}
    ]
  }'

When to use Timeline instead

A generated last frame is a target for the model, not a guarantee of pixel-exact text. If your end card carries a logo, a price or legal text, append it with Timeline 1.0. A video slot can be a still, which renders as a static hold, and the fit option controls cover, contain, stretch or blur. Slot durations must be at least 0.2 seconds, and Timeline bills $0.10 per ceil output minute, so a 12 second ad costs $0.10 to assemble.

This is also the cheaper way to test endings. Generate the body once, then build one Timeline per ending with a different still. The stored post on 10 hooks times 3 endings shows the arithmetic for that pattern.

  • Use last_frame when you want motion that resolves into a scene.
  • Use a Timeline still when text on the card must be exact.
  • Check the last second of every output either way.

Errors to expect

If you send last_frame to a model that does not list it, the API returns 400 because Sume does not silently drop unsupported fields. If both frame_images and input_references are sent, frame_images wins and the job is image-to-video. If your image URL is not publicly reachable the job fails with a message that the input media URL could not be downloaded, and the poll response says which field, not which URL.

Testing three endings without paying for three bodies

The expensive part of an ending test is the body of the ad, so avoid regenerating it. There are two sensible patterns. In the first, you generate three clips that each get a different last_frame and compare how naturally each one lands. That costs three full generations, and it tests the ending together with the motion that leads to it.

In the second pattern you generate the body once and compare the endings only. Trim the body to the length you want, then assemble three Timelines that each append a different still. Each Timeline is $0.10 for the first output minute, so three endings is $0.30 of assembly on top of the single body. If your question is which call to action wins, the second pattern isolates it, and it is the cheaper test.

Checks before you ship the clip

Pull the last frame of every output and read it at full size. The video-frames route returns exact stills at times you choose from a hosted clip, which makes this a one-call check per variant. Look for a changed logo shape, a garbled word or a color shift in the card, because a model that approximates your image can approximate its text too. If the card carries words, switch to the Timeline still.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume