Generate an image, then animate it: image to video in two API calls
Generate a still with POST /v1/images, then pass its URL as first_frame to POST /v1/videos. The two calls, the URL rule between them, and what each reserves.

Make the still with POST /v1/images, take the URL from data[0].url, and send it as a first_frame in frame_images to POST /v1/videos. That is the whole handoff: two calls, one Sume-hosted URL passed from the first to the second.
Google describes the same pattern for its own models, a fast image model feeding its Omni Flash video model. Sume's request shapes come from the Image API and Video generation docs, read 2026-09-29.
What do the two calls look like?
The image call returns 200 with the URL inside 30 seconds, or 202 with a job envelope if it is still running; in that case poll the job result for the image URL first.
curl -s -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "google/nano-banana-2", "prompt": "A red sports car on a coastal road at dusk", "aspect_ratio": "16:9"}'
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: handoff-001" \
-d '{
"model": "seedance-2",
"prompt": "Slow push in as the car pulls away",
"frame_images": [{"type": "image_url", "image_url": {"url": "IMAGE_URL_FROM_STEP_1"}, "frame_type": "first_frame"}],
"duration": 5,
"resolution": "720p",
"aspect_ratio": "16:9"
}'Do the ratios have to match?
Pick one ratio for both. The video model lists its own supported_aspect_ratios, so generate the still at a ratio the video model accepts, such as 16:9 or 9:16 for Seedance 2, rather than resizing later.
What does each step reserve?
Each call reserves its own estimate when you submit. If the wallet cannot cover it, that submit returns 402 insufficient_credits.
| Step | Model | Estimate |
|---|---|---|
| Still | google/nano-banana-2 at 1K | See the Nano Banana 2 price table |
| Clip | gemini-omni-flash-1.1, 5 s at 720p | $0.63 |
Which field is the still: frame_images or input_references?
frame_images makes the still the exact first (or last) frame, which is image-to-video. input_references treats it as guidance for reference-to-video. If you send both, frame_images wins and the request is image-to-video.
Sources
Related posts
More in Developers
- AI music generator API in Python: prompt to MP3 file
A short Python script that sends a prompt to Sume's Music Router, polls the job and saves the audio. Runs with httpx and asyncio.
- AI music negative prompt: why the API returns 400
Sume's music API rejects a non-empty negative_prompt with HTTP 400. What the error looks like and how to write exclusions in the prompt instead.
- AI music API webhook: get a callback when the track is ready
Submit a music job with mode webhook and a public HTTPS URL, verify the signed callback, and keep polling as a backup. Headers, events and retries.
- AI vendor risk assessment: questions and where to look
An AI vendor risk assessment adds training, model providers and spend to a SaaS review. The questions, and where Sume's public pages answer them.
Written by Sume