Kling 4.0's 3 to 30 seconds, split across Sume models by length
Kling 4.0 spans 3 to 30 s, and Flash 3 to 20 s. No single Sume id covers 3 to 30: Omni 3 to 10, Seedance 2.5 4 to 30, Wan 3.0 2 to 30. A lookup by seconds.

Kling 4.0 is described as producing clips of 3 to 30 seconds, with the Flash variant at up to 20 seconds, according to fal's explainer (read 2026-10-05). The full model was due in October 2026, and it is not a Sume catalog id yet. If your brief is written around 'any length from 3 to 30 seconds', here is the lookup for what you can run on Sume today.
Lookup by seconds
| Length wanted | Models that accept it | Note |
|---|---|---|
| 3 s | gemini-omni-flash-1.1, wan-3.0 | Omni 720p is $0.38; Wan 3.0 at 720p is $0.38 |
| 4 s | Omni, Seedance 2.5, Seedance 2 family, kling-3, wan-3.0, grok-imagine-video-1.5 | Grok Imagine is image-to-video only |
| 5 to 10 s | Omni (to 10 s), H3, H3 Max, Seedance, kling-3, wan-3.0 | H3 starts at 5 s |
| 11 to 15 s | Seedance, kling-3, H3, H3 Max, wan-3.0 | Omni stops at 10 s |
| 16 to 30 s | seedance-2.5, wan-3.0 | Only these two among the text-to-video models listed here; check the live catalog for image-led and recast ids |
What the table means
The edge case is 3 seconds. Past 15 seconds the text-to-video choice in this table is two ids. Read the live catalog before you commit a budget.
- The edge case is 3 seconds: only Omni and Wan 3.0 go that short.
- Past 15 seconds, among the ids in this table, you are choosing between Seedance 2.5 and Wan 3.0; the pricing is per token for one and per second for the other.
- H3 and H3 Max bottom out at 5 seconds, so a 3 or 4 second brief cannot use them.
Sources
Related posts
More in Models
- Lyria 3.5 has no edit pass: iterate a music bed for $0.125 a take
Google says Lyria 3.5 is single-turn. Iterate a bed on Sume by changing one prompt axis per take; five takes cost $0.625 and a retry needs a key.
- Every Lyria 3.5 track carries SynthID: what brands should know
Google says all Lyria 3.5 output carries a SynthID audio watermark and blocks artist voices and copyrighted lyrics. Here is what that means for brand music.
- MAI-Transcribe-2-Streaming has no server VAD: who ends the turn?
In the preview, turn_detection only accepts null, so your client sends the commit event. Sume STT has no live socket; it cuts sentences after the job finishes.
- MAI-Voice-2.1 has 23 languages, 26 locales, 28 codes: which to quote
Microsoft says 23 languages and 26 locales; OpenRouter lists 28 codes and says 30+. Quote 23 languages, and use Python to turn the 28 codes into 23.
Written by Sume