Did the model use your first frame? Check t=0 with video frames
After image-to-video, extract the frame at 0 seconds with POST /v1/video-frames and compare it to your input still. Unbilled, 1 to 24 stills per call.

To check that an image-to-video clip really starts from your still, extract the frame at 0 seconds with POST /v1/video-frames and compare it with the input. The call is unbilled, takes one workspace clip plus an at list, and returns durable image artifacts you can open next to the original.
Do this when the product or face must match exactly, such as a packshot ad. A model can honor a first frame loosely, and the check takes less time than finding the mismatch after a batch. Facts come from Sume's video frames docs and video generation docs, read 2026-10-02.
How do I send the first frame in the first place?
Use frame_images with frame_type set to first_frame. If you also send input_references, frame_images takes precedence and the request is treated as image-to-video, so a stray reference list will not override your frame. Use a public HTTPS image URL.
How do I extract the frame?
Send video_url as a media.sume.com artifact from this workspace, plus exactly one of at or fps. The Video 1.0 docs show the finished job's result carrying artifacts with a media.sume.com URL, which is what you pass. Use format: "png" for lossless inspection, and every at value must be 0 or more and less than the clip duration.
import os
import requests
r = requests.post(
"https://api.sume.com/v1/video-frames",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "first-frame-check-001",
},
json={
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"at": [0, 0.5],
"format": "png",
},
timeout=30,
)
print(r.status_code, r.json().get("request_id"))What do the limits look like?
Submit always returns 202; poll GET /v1/video-frames/{id} until resource_status is ready, then read frames[{t,url,width,height}].
| Limit | Value |
|---|---|
| Stills per call | 1 to 24 |
| Source length | Up to 300 seconds |
max_edge | 16 to 2160; omit to keep source size |
| Billing | Unbilled |
What counts as a pass?
Compare the subject, framing and colors of the t=0 still with the input. Small drift in lighting is common; a different product shape or face is a fail, so rerun with a firmer prompt or another model. Replace the demo URL with your job's artifact.
Sources
Related posts
More in Use cases
- China penalized CapCut and Dreamina over AI labels: what to check
China's CAC penalized CapCut, Maoxiang and Dreamina AI over AI labels in April 2026. What it means if your video tool serves viewers in China.
- China platform AI labels: confirmed, possible and suspected
China asks platforms to sort content as confirmed, possible or suspected AI. What triggers each label and what a video team can do about it.
- Copyright Office Part 1 on digital replicas: federal law urged
The Copyright Office AI page lists Part 1 on digital replicas (31 July 2024), which recommends federal legislation. What a voice or face project should know.
- Cost to research a competitor ad: search plus video analysis
Researching one reference video costs $0.40 on Sume: a $0.10 trending search plus a $0.30 video analysis. Ten references cost $4.00 before any generation.
Written by Sume