Sora input image must match size: the Sume frame_images rule
Sora's input image was the first frame and had to match the video size. Sume uses frame_images plus aspect_ratio, and rejects size on v1 models.

Yes, in Sora's API the input image had to match the target video's size. On Sume there is no such pairing: you send the image as a first_frame entry in frame_images, and pick the output shape with resolution and aspect_ratio. A size field returns 400 unsupported_parameter on every v1 model.
Sora facts are from OpenAI's video generation guide; Sume facts are from Video generation, both read 2026-09-30.
What did Sora require for the input image?
The guide says an input image acts as the first frame of the video, is sent as input_reference in a multipart/form-data request, and "must match the target video's resolution (size)". Its example pairs a 1280x720 size with a 720p sample image.
How does Sume take a first frame?
frame_images holds "Images for first/last frames (image-to-video)", and each entry carries a frame_type of first_frame or last_frame. The separate input_references field is for style or content guidance. If both are sent, frame_images takes precedence and the request is image-to-video, so do not send both expecting a blend.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: first-frame-001" \
-d '{
"model": "seedance-2.5",
"prompt": "Slow push-in on the product",
"resolution": "720p",
"aspect_ratio": "16:9",
"frame_images": [
{ "type": "image_url",
"image_url": { "url": "https://example.com/first.jpg" },
"frame_type": "first_frame" }
]
}'Does Sume accept a size field?
No. The docs state that "every v1 model reports supported_sizes: null, so size returns 400 unsupported_parameter", and tell you to use resolution plus aspect_ratio. Each model advertises the values it accepts in supported_resolutions and supported_aspect_ratios.
| Question | Sora guide | Sume docs |
|---|---|---|
| Field for the image | input_reference | frame_images with frame_type |
| Output shape | size such as 1280x720 | resolution + aspect_ratio |
| Must image match output? | Yes, resolution must match size | Not stated as an exact-pixel rule |
size accepted? | Yes | 400 unsupported_parameter on v1 models |
What should I do with my image?
Pick an aspect ratio from the model's supported_aspect_ratios that fits your image, check it with GET /v1/videos/models before submitting, and crop the image to that ratio yourself if it differs. Sume's docs do not describe how a mismatched first frame is fitted, so a pre-cropped image is the predictable choice. For first and last frame on a specific model, see Seedance 2.5 first and last frame.
Sources
Related posts
More in Developers
- Sora thumbnail and spritesheet variants: use Sume video frames
Sora's ?variant=thumbnail and spritesheet downloads went away with the API. On Sume, POST /v1/video-frames with at[] or fps returns durable still images.
- Sora video.completed webhook to Sume job.completed
Sora emitted video.completed and video.failed. Sume sends job.completed, job.failed and job.canceled with an x-sume-webhook-signature header. Map the handler.
- Sora Videos API replacement: seconds and size on Sume
OpenAI removed the Videos API on 2026-09-24. Map its seconds and size fields to Sume's duration, resolution and aspect_ratio; size returns 400 on Sume.
- Speech to text custom vocabulary: Sume STT takes a language hint only
Gemini 3.5 Transcribe biases up to 1,000 custom terms. Sume STT has no vocabulary field, only a language_code hint, so fix names after transcription.
Written by Sume