QC a 30-second Seedance 2.5 clip with six stills first
Extract six stills with Sume video frames at fps 0.2 and check subject, text, hands, logo and continuity before a client sees a generated 30-second clip.

Before a client sees a generated 30-second clip, extract six stills with Sume's video frames endpoint at fps: 0.2, which samples at 2.5, 7.5, 12.5, 17.5, 22.5 and 27.5 seconds, and check each against a short list. It is a spot check, not a guarantee, but it catches many visible problems cheaply. The clip must be stored on media.sume.com, so import it first.
Sampling math
The video frames docs say fps expands to mid-bin samples at 0.5/fps, 1.5/fps and so on, capped at 24 frames, with 0 < fps <= 2. At 0.2 fps over 30 seconds that gives six frames. You can instead pass 1 to 24 explicit at values. Exactly one of at or fps is allowed. Output is jpeg by default or png for lossless inspection.
| Setting | Result | Use |
|---|---|---|
| fps 0.2 | 6 stills, mid-bin | quick QC pass |
| fps 1 | 24 stills (capped) | closer review; fps 1 would sample 30 but the cap is 24, so use at[] for exact coverage |
| at [0, 14.9, 29.5] | 3 chosen instants | first, middle, last frame |
| format png | lossless | logo and text checks |
The request
Frame extraction is billed by its compute and runs as an asynchronous job, so poll it. The docs say submit always returns 202 and recommend an idempotency key on REST.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: qc-clip-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"fps": 0.2,
"format": "png"
}'Six questions per still
Ask the same short list of each frame and tick it off.
- Is it the same subject, outfit and proportions as the references?
- Are hands, fingers and faces free of obvious distortion?
- Is any text or logo legible and correct, or clearly absent?
- Does the lighting match the neighbouring stills?
- Is there anything in the frame the brief forbids?
- Does the last still end cleanly, so a caption fits?
What to do when a frame fails
If one still fails, do not guess the cause. Re-run only that idea as a short 480p probe, using the camera test grid approach, or trim the clip with video trim if the fault is at the very end. Regenerating a full 30-second clip is the expensive option, at about $17.33 for 720p on Sume's rate card.
Sources
Related posts
More in Media tools
- Swapped the voice? Re-time the visuals from word timestamps
A new voice speaks at a new pace, so cuts set for the old one drift. Take each scene's start from the new take's word timestamps and re-plan the timeline.
- Reels ad safe zone at 1080x1920 in pixels and the caption anchor
Meta says leave 14% top, 35% bottom and 6% per side clear. At 1080x1920 that is 269, 672 and 65 px; here is the anchor_ratio that keeps a caption inside.
- Reference ingest coverage: how many frames OCR read
The reference-ingest coverage block reports frames decoded, frames OCR read and the OCR rate. Use it to decide when on-screen text needs a second look.
- Reference ingest purpose: qa or remix decides who transcribes
Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.
Written by Sume