Which AI video model makes 1:1 square clips on Sume?
Kling 3.0, Wan 3.0, MiniMax H3, H3 Max and Grok Imagine list 1:1 in Sume's Videos panel; Auto does not. What to pin for square feed video.

For a 1:1 square clip on Sume, pin Kling 3.0, Wan 3.0, MiniMax H3, MiniMax H3 Max or Grok Imagine; the Auto entry in the Agents Videos panel lists only 16:9 and 9:16. If you want Auto's routing, crop a 16:9 or 9:16 result to square afterward instead.
The facts below come from the panel's per-model controls and the video generation docs.
Which models list 1:1?
The panel clamps aspect ratio to what the selected model accepts. Five of the pinned entries list square. Wan 3.0, both MiniMax entries and Grok Imagine also list 4:3 and 3:4; only the two MiniMax entries add 21:9.
| Model | Aspect ratios listed | Durations (s) | Resolutions |
|---|---|---|---|
| Auto | 16:9, 9:16 | 3, 4, 5, 6, 8, 10 | 720p, 1080p |
| Kling 3.0 | 16:9, 9:16, 1:1 | 5, 6, 8, 10, 15 | 720p, 1080p, 4K |
| Wan 3.0 | 16:9, 9:16, 1:1, 4:3, 3:4 | 2, 5, 6, 8, 10, 12, 15, 20, 30 | 480p, 720p, 1080p |
| MiniMax H3 | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 | 5, 6, 8, 10, 12, 15 | 480p, 768p |
| MiniMax H3 Max | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 | 5, 6, 8, 10, 12, 15 | 480p, 768p |
| Grok Imagine | 16:9, 9:16, 1:1, 4:3, 3:4, 2:3, 3:2 | 5, 6, 8, 10 | 480p, 720p |
What happens if I switch from Auto to a model without square?
The panel clamps; it does not error. If your current ratio is not on the new model's list it falls back to the first ratio that model lists, which is 16:9 for every entry here. That is a silent change, so check the ratio chip after switching. The settings-reset explainer covers the other controls that reset.
Which square model should I pick?
Pick by the second constraint. For a short loop under 15 seconds with a cinematic look, Kling 3.0 is the entry the panel describes as cinematic motion. For a longer clip, Wan 3.0 runs up to 30 seconds in the panel and also offers 480p. For stereo audio at native 768p, MiniMax H3 or H3 Max. The panel's per-model credit estimates differ, so compare them in the composer before a batch.
Kling's own 3.0 versus 4.0 comparison lists Kling 3.0 at 3 to 15 seconds and 720p, 1080p or 4K; Sume's panel lists 5, 6, 8, 10 and 15 seconds, so the panel is a subset of the vendor range.
- Square, 4K, up to 15 s: Kling 3.0.
- Square, long clip or 480p: Wan 3.0.
- Square with native stereo audio: MiniMax H3 or H3 Max.
- Square from a start frame, without audio: Grok Imagine (the panel lists no audio for it).
What about the API?
The API is the source of truth. Call GET /v1/videos/models and read supported_aspect_ratios for the model you intend to pin; the docs example for seedance-2 lists 1:1 among 21:9, 16:9, 4:3, 3:4 and 9:16. Then send aspect_ratio: "1:1" with your resolution. Sume's v1 models report supported_sizes: null, so size returns 400 unsupported_parameter; use resolution plus aspect_ratio.
Send an Idempotency-Key so a retry returns the original job. The finished job's usage.cost is the billable amount, so read cost from the job instead of assuming square costs the same as wide.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: square-001" \
-d '{"model":"wan-3.0","prompt":"Product on a white plinth, slow orbit","aspect_ratio":"1:1","resolution":"720p","duration":8}'What should I check before a batch?
A square result should be square on the finished job; if you pinned a model and still got a wide clip, the ratio was clamped on a model switch.
Submit one short clip first, open the finished job, and compare what you asked for with what came back. Use an Idempotency-Key on each attempt, because a replay with the same key returns the original job instead of creating and billing a second one. Only then queue the rest.
Sume reserves provider list times 1.25 when a job is submitted, and the poll response's usage.cost is the billable amount. Treat that field, not a panel estimate, as the number to budget with.
What does Sume not do?
Auto does not list square in the panel, so you cannot rely on an Auto result's shape. If a square deliverable matters, pin the model and set aspect_ratio explicitly.
Sources
Related posts
More in Media tools
- Captions with sound cues for deaf viewers: W3C checklist on Sume
W3C says captions carry speech and non-speech sound. Sume's STT burn covers speech only, so here is how to author the full cue list and burn it.
- caption_no_speech on a silent clip: burn Halloween text with cues
A silent AI clip fails video captions with caption_no_speech. Pass cues with text, start and end instead: a worked giveaway announcement on Sume at $0.20.
- Captions unreadable on busy footage: dim the clip, then burn
Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.
- Check character drift across AI shots with video_frames
Pull up to 24 evenly spaced stills from each AI clip with the unbilled video_frames route and compare faces and products before you stitch shots together.
Written by Sume