9:16 vertical AI video: which Sume models take an aspect ratio
Seedance, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash take 9:16; Grok Imagine, Genjutsu and H3 Max Recast take no aspect ratio. A per-model table.

Every text-to-video row on Sume takes 9:16: Seedance, Wan 3.0, Kling 3, MiniMax H3 and H3 Max, and Gemini Omni Flash 1.1. Three rows take no aspect_ratio at all because the output follows the input: Grok Imagine Video 1.5 (it follows the first frame), Genjutsu and H3 Max Recast (they keep the source video's framing).
Ratios per model
Only Seedance and the MiniMax rows take 21:9. Kling 3 takes just three ratios. Gemini Omni Flash takes only 16:9 and 9:16, and rejects an aspect ratio on its video edit mode.
| Model id | Accepted aspect ratios |
|---|---|
| seedance-2.5 and Seedance 2.0 rows | auto, adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| wan-3.0 | auto, adaptive, 16:9, 4:3, 1:1, 3:4, 9:16 |
| minimax-h3, minimax-h3-max | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| kling-3 | 16:9, 9:16, 1:1 |
| gemini-omni-flash-1.1 | 16:9, 9:16 |
| grok-imagine-video-1.5, genjutsu, h3-max-recast | none: do not send aspect_ratio |
What to do for a vertical ad
Send aspect_ratio: "9:16" on a text-to-video row. On Grok Imagine, crop or generate the still in 9:16 first, because the clip follows it. For a swap on an existing vertical video, the source's framing is preserved, so shoot or trim it vertical before you submit.
Rejection
Sending aspect_ratio to a row that has none is a 400. When you swap models in a script, build the request from the model's supported_aspect_ratios instead of one fixed body.
Sources
Related posts
More in Models
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume