Sume video models by aspect ratio: who takes 9:16, 21:9 and 1:1
Seedance takes 21:9 through 9:16, Kling only 16:9, 9:16 and 1:1, Gemini Omni only 16:9 and 9:16. Full matrix of Sume video models as of 2026-10-08.

Every Sume video model takes 16:9 and 9:16 except the edit and motion-transfer models, which keep the framing of the source video. Only Seedance and the MiniMax models take 21:9, only Kling and the Seedance, Wan and MiniMax models take 1:1, and Gemini Omni Flash 1.1 is limited to 16:9 and 9:16. The matrix below is from the Sume catalog, as of 2026-10-08.
Ask the live GET /v1/videos/models response for supported_aspect_ratios before you build around a ratio, because that field is the one Sume validates against.
| Model | Ratios accepted |
|---|---|
| seedance-2.5, seedance-2, -fast, -mini | auto, adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| wan-3.0 | auto, adaptive, 16:9, 4:3, 1:1, 3:4, 9:16 |
| minimax-h3, minimax-h3-max | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| kling-3 | 16:9, 9:16, 1:1 |
| gemini-omni-flash-1.1 | 16:9, 9:16 |
| grok-imagine-video-1.5 | none (follows the input image) |
| higgsfield-genjutsu, h3-max-recast | none (follow the source video) |
What the words mean
auto and adaptive let the model follow the input rather than force a shape. The Sume docs also list 3:2, 2:3 and 9:21 in the shared vocabulary of /v1/videos, but each model shows only the subset it accepts in supported_aspect_ratios.
The size parameter is not available: every v1 model reports supported_sizes: null, so size returns 400 unsupported_parameter. Use resolution plus aspect_ratio.
Pixel sizes for a ratio
Seedance prices on pixels, so the actual frame matters. In the Sume pricing tables, 720p in 9:16 is 720 x 1280, in 16:9 is 1280 x 720, in 1:1 is 960 x 960 and in 21:9 is 1470 x 630. At 1080p, 9:16 is 1080 x 1920 and 1:1 is 1440 x 1440.
| Resolution | 16:9 | 9:16 | 1:1 | 21:9 |
|---|---|---|---|---|
| 480p | 864 x 496 | 496 x 864 | 640 x 640 | 992 x 432 |
| 720p | 1280 x 720 | 720 x 1280 | 960 x 960 | 1470 x 630 |
| 1080p | 1920 x 1080 | 1080 x 1920 | 1440 x 1440 | 2205 x 945 |
Examples
A 9:16 clip for short-form feeds can come from every generation model in the table. At 720p and 5 seconds, that is $0.625 on Wan 3.0, $0.625 on Gemini Omni, $0.375 on MiniMax H3 at 768p, and $2.889 on Seedance 2.5. A 21:9 clip needs Seedance (from about $0.95 for 5 seconds at 720p on Mini) or MiniMax ($0.375 on H3 at 768p).
If your pipeline switches ratios per channel, pick a model that accepts all of them. Seedance and the MiniMax models cover the widest range; Gemini Omni and Kling cover the narrowest.
Picking a ratio by channel
The Sume Auto model offers 3 to 10 second clips at 16:9 or 9:16 in its create controls.
- Vertical feeds: 9:16, accepted by every generation model in the table.
- Square placements: 1:1 works on Seedance, Wan 3.0, MiniMax and Kling, not on Gemini Omni.
- Cinematic bars: 21:9 on Seedance or MiniMax only.
- Reframing existing footage: use a model that follows the source, such as the edit and recast models.
Sources
Related posts
More in Developers
- AI video API with audio reference input: which Sume models accept it
Seedance (all four), Wan 3.0, MiniMax H3 and H3 Max accept audio references on Sume. Kling, Gemini Omni, Grok Imagine and the motion models do not.
- Video frames caps at 24 stills: a 15 s hook needs fps 1.6, not 2
Video frames returns at most 24 stills and fps up to 2. At 2 fps only 12 s are covered, so a 15 s hook needs fps 1.6; the Python below checks any length.
- One env line picks the video model: $1.25 vs $5.78 per 10 s clip
Read the Sume video model id from one env var. For a 10 s 720p 9:16 clip, gemini-omni-flash-1.1 bills $1.25 and seedance-2.5 bills $5.78.
- Clip-type route table in Python: Sora use cases to Sume models
A 30-line Python route table maps ads, demos and teasers to a Sume model, size and length, refuses a length the model cannot make, and prints the dollar cost.
Written by Sume