CTA end card for an AI video ad: use last_frame on Sume
End an ad clip on your CTA card by sending it as a last_frame image on /v1/videos. Which models accept it, which do not, and the Timeline alternative.

To make a video ad end on your call-to-action card, send the card as a frame_images entry with frame_type last_frame in a POST /v1/videos body. Sume accepts a first frame, a last frame or both, but only on models whose catalog entry lists last_frame in supported_frame_images. If you need the card to appear exactly, unchanged and for a fixed number of seconds, use Timeline 1.0 instead and append the still as a static hold.
Which models take a last frame
The Video Router catalog marks an end-frame capability per model, and the /v1/videos descriptor turns that into supported_frame_images. The table is taken from the capability records in the Sume repository and should be confirmed against GET /v1/videos/models for the model you use.
| Model id | Last frame | Duration range (s) |
|---|---|---|
| seedance-2.5 | Yes | 4 to 30 |
| seedance-2, seedance-2-mini, seedance-2-fast | Yes | 4 to 15 |
| kling-3 | Yes | 4 to 15 |
| wan-3.0 | Yes | 2 to 30 |
| minimax-h3, minimax-h3-max | Yes | 5 to 15 |
| gemini-omni-flash-1.1 | Yes | 3 to 10 |
| grok-imagine-video-1.5 | No | 4 to 15 |
| h3-max-recast | No (edit-only model) | 5 to 30 source |
The request
Each frame is an object with type image_url, an image_url object holding a public URL, and a frame_type. If you also want the clip to start from a product still, add a second entry with frame_type first_frame. The API infers image-to-video from the presence of frame_images, so no mode field is needed.
The model then generates motion that ends on, or near, your last frame. The result is a clip that arrives at the card, not a clip with a card pasted on. You should still expect small differences in how closely the final frame matches your image, so check the last second of each output before you ship it.
curl -sS -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mug-endcard-v1" \
-d '{
"model": "seedance-2",
"prompt": "Hand sets a mug down, camera pushes in",
"duration": 6,
"aspect_ratio": "9:16",
"frame_images": [
{"type":"image_url","image_url":{"url":"https://example.com/end-card.png"},"frame_type":"last_frame"}
]
}'When to use Timeline instead
A generated last frame is a target for the model, not a guarantee of pixel-exact text. If your end card carries a logo, a price or legal text, append it with Timeline 1.0. A video slot can be a still, which renders as a static hold, and the fit option controls cover, contain, stretch or blur. Slot durations must be at least 0.2 seconds, and Timeline bills $0.10 per ceil output minute, so a 12 second ad costs $0.10 to assemble.
This is also the cheaper way to test endings. Generate the body once, then build one Timeline per ending with a different still. The stored post on 10 hooks times 3 endings shows the arithmetic for that pattern.
- Use last_frame when you want motion that resolves into a scene.
- Use a Timeline still when text on the card must be exact.
- Check the last second of every output either way.
Errors to expect
If you send last_frame to a model that does not list it, the API returns 400 because Sume does not silently drop unsupported fields. If both frame_images and input_references are sent, frame_images wins and the job is image-to-video. If your image URL is not publicly reachable the job fails with a message that the input media URL could not be downloaded, and the poll response says which field, not which URL.
Testing three endings without paying for three bodies
The expensive part of an ending test is the body of the ad, so avoid regenerating it. There are two sensible patterns. In the first, you generate three clips that each get a different last_frame and compare how naturally each one lands. That costs three full generations, and it tests the ending together with the motion that leads to it.
In the second pattern you generate the body once and compare the endings only. Trim the body to the length you want, then assemble three Timelines that each append a different still. Each Timeline is $0.10 for the first output minute, so three endings is $0.30 of assembly on top of the single body. If your question is which call to action wins, the second pattern isolates it, and it is the cheaper test.
Checks before you ship the clip
Pull the last frame of every output and read it at full size. The video-frames route returns exact stills at times you choose from a hosted clip, which makes this a one-call check per variant. Look for a changed logo shape, a garbled word or a color shift in the card, because a model that approximates your image can approximate its text too. If the card carries words, switch to the Timeline still.
Sources
Related posts
More in Developers
- How do I narrate a DIY tutorial step by step with a TTS API?
Narrate an 8-step DIY tutorial with one TTS job per step: 1,570 characters, $0.10 on Sume. Why per-step jobs make a fixed step a 1-cent redo.
- Do I pay for a failed AI avatar video job? Refunds on Sume
Sume reserves the avatar video price at submit, captures it on completion, and releases or refunds it where a job fails. What it means for retries.
- Does PNG, JPEG or WebP change the price of an AI image on Sume?
No. On Sume's Image API, output_format picks the file type, not the price: per-image cards and GPT Image 2.5 token math ignore it. Which models list which.
- duration vs duration_seconds on each Sume video route
/v1/videos takes duration; motion control and lip-sync take duration_seconds; recast and edit read the source clip. One table of what each does.
Written by Sume