AI video models with 4:3 and 3:4 aspect ratios on Sume
Seedance 2.5 and 2.0, Wan 3.0, MiniMax H3 and H3 Max list 4:3 and 3:4 on Sume; Kling 3.0 and Gemini Omni Flash 1.1 do not. Prices for the per-second rows.

4:3 and 3:4 are the classic photo and tablet frames, and they sit between 1:1 and the wide and tall video ratios. Not every Sume video row lists them. The catalog is explicit: a ratio a model does not advertise is a 400, so the question is which rows list both.
Seedance 2.5, the Seedance 2.0 family, Wan 3.0, MiniMax H3 and MiniMax H3 Max all list 4:3 and 3:4, along with 16:9, 9:16 and 1:1 and the adaptive option. Kling Video v3 Pro lists only 16:9, 9:16 and 1:1, and Gemini Omni Flash 1.1 only 16:9 and 9:16.
What does each 4:3-capable row cost?
Aspect ratio does not change the price on a per-second row; the rate is keyed by resolution. So the costs are the same for 4:3 as for 16:9. The table shows the billable rate for the per-second rows.
| Row | Resolution | Billable per second |
|---|---|---|
wan-3.0 | 480p / 720p / 1080p | $0.0625 / $0.125 / $0.25 |
minimax-h3 | 480p / 768p | $0.0625 / $0.075 |
minimax-h3-max | 480p / 768p / 1080p | $0.0625 / $0.10 / $0.20 |
| Seedance rows | 480p / 720p / 1080p | per video token |
Which should I pick?
For a 30-second 3:4 clip, Seedance 2.5 and Wan 3.0 are the only options, since the MiniMax rows stop at 15 seconds. For the cheapest 4:3, Wan 3.0 and the MiniMax rows tie at $0.0625 per second at 480p. For the top resolution at 3:4, H3 Max reaches 1080p at $0.20 per second, but note the catalog says 1080p there is latent refinement from native 768p, not a native 1080p render.
If you need audio references or reference images, all of these accept them; Kling does not.
How do I request it?
Send aspect_ratio: "3:4" or "4:3" on POST /v1/videos. Use adaptive if you pass a first frame and want the output to follow its shape. Do not send size; Sume's catalog does not accept pixel sizes on these rows.
Sources
Related posts
More in Models
- AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
- Arabic text to speech API: send language ar to Sume TTS
Arabic is on Cartesia's Sonic 3.6, 3.5 and 3 but never on Sonic 2. How to send ar through Sume TTS, which voice id to use, and what an Arabic script costs.
- AudioCraft weights are CC-BY-NC: which models that covers
AudioCraft's code is MIT but its README puts the model weights under CC-BY-NC 4.0, including MusicGen and AudioGen. What that means for ad and video audio.
Written by Sume