Sume auto or a pinned video model: what changes past 10 seconds
Auto video defaults to 720p and 8 seconds with 3 to 10 second clips. For 12, 15 or 30 seconds, 1080p or 21:9, pin a model id. How to decide on Sume.

Use sume/auto when a 3 to 10 second clip at 16:9 or 9:16 is enough, and pin a model id when you need more. The docs describe Auto's create controls as defaulting to 720p and 8 seconds, with 3 to 10 second clips at 16:9 or 9:16. A 15 or 30 second clip, a 21:9 frame or a specific reference setup is a reason to pin.
What Auto gives up
Auto's responses echo sume/auto and Sume does not disclose which family ran, so you cannot build a workflow that depends on one model's look. That is the point for a quick clip and a drawback for a series that needs consistency.
When to pin
The rest of the catalog is explicit pass-through: you choose the id, Sume validates the request against that model's envelope, and the receipt names the model.
| You need | Pin |
|---|---|
| 30 seconds in one clip | seedance-2.5 or wan-3.0 |
| 2-second clip | wan-3.0 |
| 4K output | gemini-omni-flash-1.1 |
| Edit an existing video with a prompt | gemini-omni-flash-1.1 with video_url |
| Reference audio | seedance-2.5, wan-3.0 or minimax-h3 |
| Always-on stereo audio, 1080p | minimax-h3-max |
Migrating
Auto is a model value on POST /v1/videos, so moving from Auto to a pinned model is one field. Keep your Idempotency-Key scheme per request so a retried submit returns the original job.
Sources
Related posts
More in Models
- 9:16 vertical AI video: which Sume models take an aspect ratio
Seedance, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash take 9:16; Grok Imagine, Genjutsu and H3 Max Recast take no aspect ratio. A per-model table.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
Written by Sume