Preview a 3-minute Short with 24 stills, one every 7.5 seconds

Video frames returns up to 24 stills per call, which spaces a 180-second Short at 7.5 seconds. Build the at[] list in Python and submit it to /v1/video-frames.

5 min readSume
All posts

Short answer

A Short can run up to 3 minutes (read 2026-10-05), and POST /v1/video-frames takes at most 24 timestamps per call. 180 seconds divided by 24 is 7.5 seconds, so one call can show every 7.5-second window of a full-length Short. Put each timestamp in the middle of its window: 3.75, 11.25, and so on to 176.25.

Why a contact sheet beats scrubbing

A long vertical Short has a lot of places to go wrong: a clip that did not match its neighbor, a caption that cut off, a black frame at an edit point. Twenty-four stills in one look is faster than watching three minutes. Video frames decodes the accurate frame at each instant at the source size, with an optional max_edge clamp from 16 to 2160 and jpeg or png output.

The route accepts a source up to 300 seconds, so it covers the full Shorts length. It always returns 202 with a job, so you poll the job or the video_frames resource.

Spacing options

All rows use the 24-still cap per call.

Still spacing for 24 frames, read 2026-10-05
Clip lengthSeconds per stillFirst mid-window timestampLast mid-window timestamp
60 s2.51.2558.75
90 s3.751.87588.125
180 s (Shorts maximum)7.53.75176.25
300 s (video frames source cap)12.56.25293.75

Submit the call

The script builds the list and sends it. It asks for a 540-pixel long edge so the stills are easy to tile, and it sends an Idempotency-Key so a retry does not queue a second extract. Replace the URL with your own media.sume.com artifact.

import json, os, urllib.request

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("set SUME_API_KEY")
length, n = 180, 24
at = [round((i + 0.5) * length / n, 3) for i in range(n)]
body = {"video_url": "https://media.sume.com/artifacts/artf_demo/short.mp4",
        "at": at, "max_edge": 540}
req = urllib.request.Request(
    "https://api.sume.com/v1/video-frames",
    data=json.dumps(body).encode(),
    headers={"Authorization": "Bearer " + key,
             "Content-Type": "application/json",
             "Idempotency-Key": "preview-short-180"},
)
print(json.load(urllib.request.urlopen(req))["request_id"], at[:3], at[-1])

What to read in the stills

  • Frame 1 and frame 24: the opening and closing beats. The last frame, 176.25 s, still sits inside the final window, not at the very end.
  • Frames that look identical in a row: a clip held too long or a timeline slot that padded a short source.
  • Text near the edges: whether burned-in captions survive in a 540-pixel view.
  • Anything black: if you need the exact instant, request a denser at[] around that window.

Cost and limits

Video frames is billed by its own Modal compute rather than a flat per-job price, and the bill never exceeds the hold taken at submit. The docs I read give no flat rate, so check GET /v1/catalog before a large batch. For a coarser look at one clip, the default video inspect call returns 8 evenly spaced stills with a 768-pixel long edge and also gives you the probe, which is useful for checking duration against the 180-second ceiling before you spend the 24-still call.

If the source is longer than 300 seconds, trim the part you want first, then run this call on the trimmed clip.

For a quick sanity pass, sort the stills into a grid with any image tool and look for the three things that matter in a long Short: the hook at the start, a visible change every few windows, and a clean ending. If you find a problem at, say, the 9th still, the window is 63 to 70.5 seconds, which tells you where to open the edit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume