Read a competitor ad before you remix it: 24 frames per call
Sume video-frames pulls up to 24 stills from a clip you own, at times you name, so you can study a reference ad structure before a remix. Limits and errors.

To study a reference ad before a remix, import the clip into your workspace and call Sume's video frames endpoint with up to 24 timestamps (at[]) or a sample rate (fps up to 2). It returns durable image artifacts at source size, from a source up to 300 seconds. Use only clips you have the right to analyze or recreate, such as your own earlier ads, licensed footage, or material you have permission for.
The goal is to understand structure: where the hook sits, when the product shows, how long each shot holds and what text appears. You then brief your own version in your own words and your own footage.
Pick the instants
A 30 second reference with six beats needs six or seven stills, not 24. Choose at[] from your own viewing: the first frame, each cut, the product reveal and the end card. Each value must satisfy 0 <= t < duration, or the worker fails with frame_time_out_of_range and tells you the probed duration.
| Field | Limit |
|---|---|
| at[] | 1 to 24 values |
| fps | 0 < fps <= 2, capped at 24 frames |
| Source length | 300 s maximum |
| format | jpeg (default) or png |
| max_edge | 16 to 2160 |
| Submit status | Always 202; poll the resource |
Call it
Send exactly one of at[] or fps, never both. The clip must already be a media.sume.com artifact, because the API does not fetch from the open internet. Add an Idempotency-Key so a retry does not queue a second extract.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ref-ad-frames-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"at": [0, 2.5, 6, 11, 14.5],
"format": "png"
}'Read the result
Poll GET /v1/video-frames/:id. When resource_status is ready, frames[] holds t, url, width and height for each still, and source_duration_seconds is the probed length. If one instant fails, that frame has url: null; the job itself does not fail, so check each entry. This job uses no admission seat and is billed by its Modal compute, never more than the hold.
Do not use it where you need a transcript or an overview; the docs point to video inspect for a probe, eight mid-bin stills and optional speech-to-text, and to video trim for a range cut.
| Question | Tool |
|---|---|
| What is at exactly 6.0 s? | video frames |
| What is in this clip overall? | video inspect |
| Give me seconds 2 to 8 as a clip | video trim |
Turn stills into a brief
Lay the stills in a row, label each with its time, and write one line per shot: what the viewer sees, why it works and what you will shoot instead. Keep the structure and rewrite the content, because a remix that copies a competitor's footage, words or branding is a legal and brand risk, not a creative strategy.
Then build your version with your own product clips and a fresh script, and re-check it with the same frame call at the same beats.
- Own or licensed clips only.
- Six to eight stills usually tell the structure.
- Brief in your words, shoot your product.
Related posts
More in Use cases
- Real estate agent intro clip: headshot, logo and listing photos
Send a headshot, your logo and listing photos as reference_image_urls to Gemini Omni Flash 1.1 and address each as <IMAGE_REF_n> for a 3 to 10 second intro.
- Listing clip from photos with Seedance 2.5, and what to disclose
Nine listing photos can feed a Seedance 2.5 reference-to-video request on Sume. Price a 30 s clip, plan the camera path, label it AI-generated.
- Real estate price-reduction clip from one listing photo: $0.83 each
A 5-second vertical price-drop clip from the hero photo with the new price burned in as cues: $0.825 per listing, $9.90 for 12, on Sume from docs.
- A listing walkthrough in English and Spanish from one silent video
Narrate one silent walkthrough twice, with two text-to-speech tracks at $0.0475 per 1,000 characters, then join each with captions. About $0.35 a language.
Written by Sume