Proof frame for a 1080x1920 export: video frames PNG at 1920
Pull a lossless 1080x1920 still from your vertical export with video frames (format png, max_edge 1920) to check captions and crop before you upload.

To proof a 1080x1920 export before upload, extract one or two lossless stills with video frames: at: [1, 12], format: "png". Frames keep the source size when you leave max_edge out, so a 1080x1920 export comes back at 1080x1920. Setting max_edge: 1920 is the same size for a vertical clip and makes the intent explicit. The job returns durable image artifacts, never modifies the source, and fails a single frame softly (a null url) without failing the job.
What to look at in the still
A proof frame answers questions a probe cannot: whether the crop cut off a face, whether burned-in captions are inside the frame, and whether text is legible at full size.
| Field | Value | Why |
|---|---|---|
| at | [1, 12] (1 to 24 values) | Pick one frame with a caption and one without |
| format | png | Lossless; jpeg is the default |
| max_edge | omit or 1920 (range 16 to 2160) | Keeps 1080x1920 |
| source length | 300 s or less | Frames limit |
The request
Submit returns 202 always; frames does not support mode: "sync". Poll GET /v1/video-frames/:id until resource_status is ready, then read frames[{t,url,width,height}]. Billing is by the worker's compute, with the reserve capped at the hold, so a two-frame proof is a small job.
Import the clip first with POST /v1/media-imports, and send an Idempotency-Key on writes.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: proof-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/export.mp4",
"at": [1, 12],
"format": "png",
"max_edge": 1920
}'Common failures
frame_time_out_of_range: anatvalue is at or past the duration; the error reports the probed duration.duration_out_of_range: the source is longer than 300 seconds; cut it with video trim first.- A 400 for sending
at[]andfpstogether, or more than 24 values. - For probe facts or a transcript instead, use video inspect.
A three-frame routine
For a 42.5 second export, three frames cover most failures: one near the start, one in the middle where captions are densest, and one just before the end. The last frame must be before the duration, so use the duration minus a small margin, not the duration itself.
| at (s) | What it shows |
|---|---|
| 0.5 | The opening frame and crop |
| 21 | Captions at their densest |
| 42.4 | The end card; 42.5 itself would fail with frame_time_out_of_range |
What to compare
Open each still at full size beside the platform's overlay. Look for three things: any caption word that touches a screen edge, any face or product that the crop clipped, and any text inside the picture that is too small to read on a phone. If you find one, fix it at the earliest step that caused it: the crop numbers in video filter, the placement in captions, or the fit in Timeline. Then repeat the proof on the new file.
Because frames are durable media.sume.com artifacts, you can keep a proof beside each final file as a record of what you checked. A null frame url means that one instant could not be extracted; the job still succeeds, so inspect the result before assuming all frames exist.
The same routine works for any vertical export from this series, whether it came from a crop, a trim or a Timeline render. It is cheap to run, it does not touch the file, and it finds the mistakes that matter most before the audience does: a clipped face, a caption that runs off the edge, or a frame that is the wrong shape.
Sources
Related posts
More in Media tools
- Silent reference clip: reference_ingest's -60 LUFS gate and STT
reference_ingest marks a clip audio.silent at -60 LUFS or below and skips STT, so a silent reference costs no transcript. What to do next: plan new music.
- reference-ingest source_no_video_stream: audio-only file as input
source_no_video_stream means the reference file has no video track, such as an m4a. Send a clip with picture, or use audio detach to work with the sound.
- reference_ingest_stt_required: speech.language_code without STT
speech.language_code and duration_seconds only apply when the read transcribes. Set speech.allow_billed_stt true, use reference_remix, or drop them.
- Reference ingest text_tracks is not caption coverage: 5 sampled frames
text_tracks come from OCR on deduplicated frame states, at most five frames, not every caption. To know if a Short is captioned, check frames yourself.
Written by Sume