Product photo to a 9:16 Reel clip in Python: 6 seconds for $0.75
POST /v1/videos with one product photo as frame_images first_frame on gemini-omni-flash-1.1: 6 seconds at 720p costs $0.75. Python polls and saves the MP4.

To turn a product photo into a 9:16 Reel clip, send the photo as frame_images with frame_type: "first_frame" to POST /v1/videos, set aspect_ratio to 9:16, and poll the job. On gemini-omni-flash-1.1 at 720p, a 6-second clip costs $0.75 (6 x $0.125). The code below runs as written once SUME_API_KEY is set and the photo URL is a public HTTPS address.
Image-to-video matters more this autumn because video models are moving. The ElevenLabs changelog (read 2026-10-05) records that OpenAI discontinued Sora 2 and Sora 2 Pro on September 24 and that ByteDance deprecated Seedance 1.5 Pro, retiring November 11. A product-shot pipeline that hard-coded one of those ids needs a new model id, and a catalog lookup is a safer place to take it from.
Why this model and these fields
The Video Router table lists gemini-omni-flash-1.1 with durations from 3 to 10 seconds, resolutions of 360p, 720p, 1080p and 4K, and image-to-video with an optional end frame. It accepts 16:9 and 9:16 only, which suits a vertical Reel. The list price at 720p is $0.10 per second, and Sume bills list times 1.25, so one second is $0.125.
On /v1/videos the API infers the mode from the fields. If frame_images is present, the job is image-to-video. If input_references is present, it is reference-to-video. If neither is present, it is text-to-video. A pinned opening frame belongs in frame_images with frame_type: "first_frame", so the photo becomes the first frame of the clip.
| Field | Value in this example | Why |
|---|---|---|
| model | gemini-omni-flash-1.1 | 3 to 10 s, 9:16, image-to-video |
| frame_images[0].frame_type | first_frame | Photo is the opening frame |
| aspect_ratio | 9:16 | Vertical Reel; this model takes 16:9 or 9:16 |
| resolution | 720p | $0.125 per second after the 1.25 factor |
| duration | 6 | 6 x $0.125 = $0.75 |
The script
The request carries an Idempotency-Key, so a retry after a network error does not pay twice. The poll reads polling_url from the submit response and stops on a terminal status. Statuses on this surface are pending, in_progress, completed, failed and cancelled.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
"model": "gemini-omni-flash-1.1",
"prompt": "Slow push-in on the product, soft window light.",
"duration": 6, "resolution": "720p", "aspect_ratio": "9:16",
"frame_images": [{
"type": "image_url",
"image_url": {"url": "https://example.com/product.jpg"},
"frame_type": "first_frame",
}],
}
r = requests.post("https://api.sume.com/v1/videos", json=body,
headers={**H, "Idempotency-Key": "reel-shot-001"}, timeout=30)
r.raise_for_status()
job = r.json()
while job["status"] in ("pending", "in_progress"):
time.sleep(10)
job = requests.get(job["polling_url"], headers=H, timeout=30).json()
print(job["status"])
if job["status"] == "completed":
url = f"https://api.sume.com/v1/videos/{job['id']}/content?index=0"
open("shot.mp4", "wb").write(requests.get(url, headers=H, timeout=120).content)Cost and what to check
One shot is $0.75. Four products are $3.00, and a first take that you reject costs the same again. Read the live rate from GET /v1/videos/models before you plan a batch, because the numbers here are the rates in the repository on 2026-10-05.
Check the first second of the result. The photo is the opening frame, but the model moves from it, so the product can drift or change over the clip. If the product must stay exact, keep the clip short and compare the last frame to your photo. A failed job returns 409 job_failed from the content route, which is not retryable, so stop polling when the status is failed instead of looping on the content URL.
Prepare the photo too. It has to be a public HTTPS URL that Sume can fetch, so a local path or a signed private link will not work. Use a clean, well-lit shot with the product centered and some empty space around it, since the model animates from that frame. If you want the clip to end on a specific frame, the frame_images array also accepts a second image with frame_type: "last_frame" on models that support an end frame. A reasonable Reel workflow is one first-frame job per product, then a manual review of the six seconds before you spend on captions or music.
Sources
Related posts
More in Developers
- Prometheus histogram buckets for Sume video jobs, up to 20 minutes
Pick histogram buckets for submit-to-terminal time on 30-second Seedance 2.5 and other Sume video jobs, so the 20-minute SDK deadline is the last bucket.
- Promise.allSettled for a wave of Sume jobs: keep the partial wins
Promise.all throws away nine good Sume videos when one job fails. Use Promise.allSettled over a wave sized to your plan's concurrency, in 30 lines of Node.
- provider_submission_failed 502: Sume could not start the job, retry
provider_submission_failed is a 502 meaning the job never started. By default it is retryable with retry_after_seconds 30 and next_action retry_later.
- Pydantic v2 models for the Sume /v1/videos poll response
Typed Pydantic v2 models for POST and GET /v1/videos: five status literals, an optional error string, a url list, and a guard that stops a bad status early.
Written by Sume