Check character drift across AI shots with video_frames
Pull up to 24 evenly spaced stills from each AI clip with the unbilled video_frames route and compare faces and products before you stitch shots together.

To check whether a character or product drifted inside a generated shot, extract evenly spaced stills with POST /v1/video-frames using fps and look at them side by side. The route is unbilled, returns durable images at source size, and takes at most 24 frames per call.
Sume does not score likeness or compare faces for you. The docs describe frame extraction only, so the judgement is yours (or your own reviewer's). What the route gives you is cheap, repeatable evidence for that judgement, taken before you commit a shot to a timeline.
Why sample frames instead of scrubbing the clip?
Drift can appear anywhere in a clip, and the first frame may be your own input image, so one still is not enough. A grid of stills at fixed intervals makes the change visible in one glance: the jacket color, the logo position, a ring that appears on the other hand.
fps is the sampling rate and must satisfy 0 < fps <= 2. The route expands it into mid-bin samples at 0.5/fps, 1.5/fps, 2.5/fps and so on, capped at 24 frames. For a 12-second clip, fps: 1 gives frames at 0.5, 1.5 ... 11.5 seconds.
How do I request the sheet?
Required: video_url on your workspace's media.sume.com (generated clips already are) and exactly one of at[] or fps. Optional: format (jpeg default or png) and max_edge between 16 and 2160. Submit always returns 202; poll GET /v1/video-frames/:id. When resource_status is ready, frames[] holds t, url, width and height for each still. One instant that fails extraction comes back with a null url and does not fail the job.
Send an Idempotency-Key on REST so a retry does not queue a second extract. The script below saves the timestamps next to the URLs, which is all you need to build a grid in whatever viewer you use.
import os, time, requests
BASE = "https://api.sume.com/v1/video-frames"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
clip = "https://media.sume.com/artifacts/artf_demo/shot-03.mp4"
r = requests.post(BASE, headers={**H, "Idempotency-Key": "sheet-shot-03"},
json={"video_url": clip, "fps": 1, "format": "png", "max_edge": 768})
r.raise_for_status()
rid = r.json()["request_id"]
deadline = time.time() + 120
while time.time() < deadline:
body = requests.get(f"{BASE}/{rid}", headers=H).json()
body = body.get("video_frames", body)
if body.get("resource_status") == "ready":
for f in body["frames"]:
print(f"{f['t']:5.1f}s {f['url']}")
break
time.sleep(2)
else:
raise TimeoutError("frames not ready")What limits apply to a contact sheet?
A warning is possible: low_confidence_long_video appears when the source is longer than 90 seconds. It does not fail the job, and AI shots are rarely that long.
| Limit | Value | Effect on review |
|---|---|---|
| Frames per call | 24 maximum | One call covers a 12-second clip at fps 2 |
| fps | Above 0, at most 2 | Dense sampling needs several calls with at[] |
| Source length | 300 seconds maximum | Shot clips are well inside this |
| max_edge | 16 to 2160 | Omit it to keep source size for pixel-level checks |
| Format | jpeg or png | png is lossless, better for judging fine product detail |
| Billing | Unbilled | Review every shot, not just the suspicious ones |
What do I compare, and what do I do about drift?
Pick two or three things per shot and write them down before you look: face shape and hair, one garment, the product label. Compare each sheet to your approved still, not to the previous shot, because drift accumulates across a chain.
If a shot drifts, the fix is upstream, not in the timeline. Regenerate that one shot with the approved still as the first_frame (or as an input_references entry on models that take references), keeping the prompt short on appearance and long on motion. Because every shot is its own job, one regeneration does not touch the others. The related post on consistent characters across shots covers the prompting side.
For whole-clip evidence in one call, video inspect returns a probe plus eight sampled stills at a 768 default edge. Use video_frames when you want a specific count, a specific timestamp, or full-size images.
Sources
Related posts
More in Media tools
- Check an AI take says your script: hypit align and unmatched words
hypit align pairs each script token with the transcript words of a generated take, and lists unmatched words. What it measures and what it does not do.
- Join voiceover takes into one gapless track with Timeline audio
Concatenate up to 20 Sume-hosted voice takes with Timeline audio, with no seam silence and no re-synthesis, and re-base video starts from the returned offsets.
- Crop a 16:9 video to a centered 9:16 strip: the crop fractions
For a centered 9:16 crop of a 16:9 video, send crop x 0.3418, y 0, width 0.3164, height 1 to Sume video-filter. 1:1 and 4:5 values are in the table.
- Cyber Monday ad in 9:16, 1:1 and 16:9 from one clip with Timeline
Resize one offer video to vertical, square and landscape with Timeline 1.0 output width and height, fit blur or cover, for $0.30 across three renders.
Written by Sume