Transloadit /video/thumbs offsets vs Sume video frames at[]
Transloadit /video/thumbs takes up to 999 thumbnails by offsets or count. Sume video frames returns at most 24 stills per call, by at[] seconds or fps up to 2.
Transloadit's /video/thumbs robot can return up to 999 thumbnails from one video, picked either by a count or by an offsets array of seconds. Sume's video frames route is smaller by design: 1 to 24 stills per call, chosen with at[] or an fps of at most 2, returned as durable image artifacts.
Transloadit's side comes from its /video/thumbs documentation, Sume's from the video frames doc, both read on 2026-10-03.
What do Transloadit's offsets and count do?
count defaults to 8 and has a stated maximum of 999. offsets is an array of seconds such as [ 2, 45, 120 ], and the page says it supports percentages and decimals for milliseconds. If both are given, offsets wins and cannot be combined with count.
The output defaults are format jpeg (jpeg, jpg or png accepted), the original video width and height, and a resize_strategy of pad.
| Parameter | Default | Notes |
|---|---|---|
| count | 8 | Maximum 999 |
| offsets | [] | Seconds, percentages, decimals; overrides count |
| format | jpeg | jpeg, jpg or png |
| width / height | Original | String or number |
| resize_strategy | pad | crop, fit, fillcrop, min_fit, pad, stretch |
How does Sume video frames select stills?
Send video_url plus exactly one of at[] or fps. at[] takes 1 to 24 explicit seconds, each at least 0 and below the clip duration. fps must be above 0 and at most 2; Sume expands it to mid-bin samples (0.5/fps, 1.5/fps, and so on) and caps the result at 24 frames.
format is jpeg (default) or png, and max_edge between 16 and 2160 clamps the long edge. Leave it out and you keep the source frame size. Source clips must be 300 seconds or shorter.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: thumbs-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"at": [2, 45, 120],
"format": "png",
"max_edge": 1280
}'What do you lose when moving from Transloadit to Sume?
Count is the obvious loss. A Transloadit job can return hundreds of thumbnails for a scrub bar, and Sume tops out at 24 per call. For a longer sprite you would make several calls with different at[] lists. The 300 second source limit applies to every call, so for longer videos video inspect (source up to 1800 seconds, 24 stills) and a trim first are the options.
Percent offsets are the other loss. Sume wants seconds, so to get 10 percent steps you read the clip length first. Its worker reports source_duration_seconds on a finished video frames resource, and a probe-only video inspect (frames: false) returns probe facts without stills.
Pad also has no direct match: Sume frames come out at source size, or clamped by max_edge, with no pad or crop strategy.
How would you build a 100-thumbnail scrub bar on Sume?
You would not do it in one call. Split the source into windows with video trim, then request up to 24 stills per window. A 300 second clip at one still every 12 seconds is 25 stills, so two calls cover it. Each call is async, so submit them with distinct Idempotency-Key values and poll each resource.
Be honest about the trade-off: Transloadit is built for dense thumbnail sets, and Sume video frames is built for choosing a handful of exact frames, such as a hook frame, a cover candidate or a QA check. If your product needs hundreds of thumbnails per upload, the Transloadit robot fits better.
What do you gain?
Sume returns durable media.sume.com artifacts with t, url, width and height per frame. One frame that fails to extract comes back with a null url and does not fail the job. A frame_time_out_of_range error names the probed duration, which makes the fix obvious.
Submit is always 202, so poll GET /v1/video-frames/:id until resource_status is ready. The job is billed by its own Modal compute, and the jobs and results page covers polling and webhooks.
Sources
Related posts
More in Media tools
- Trim a clip to 10 seconds before a Gemini Omni edit
Google caps Omni edit and extend inputs at 10 seconds. Cut longer footage first with Sume's video-trim endpoint, a flat $0.02 per job.
- Turn a ChatGPT try-on image into a video with Seedance 2.5
Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.
- Turn a hum into music with AI: what Sume takes as input
Stability says hum-to-steer is coming. Sume's music API takes text and one optional image, not audio. Here is how to describe a hummed tune in a prompt.
- Upscale an old Sora download: the limits of Sume's video upscaler
Sora files you saved can be upscaled. Sume's video-upscale model takes a scale from 1.1 to 4, up to 30 seconds, and a fast, standard or pro tier.
Written by Sume